Staff Software Engineer, Kubernetes Platform

Anthropic • $320K — $405K *

San Francisco, CA 94112In-Person

Information Technology

8 - 10 years of experience

3 weeks ago

By clicking Apply, I agree with Ladders' Terms of Use and Privacy Policy

Job Overview by Ladders

Qualifications

5-7 years of software engineering experience with distributed systems
Proficiency in systems programming languages such as Go, Python, Rust, or C++
Extensive hands-on experience with Kubernetes internals, especially in scheduling and control plane operations
Strong debugging skills for complex issues in distributed systems
Experience designing reliable systems that support multiple engineering teams
Excellent communication skills for stakeholder interactions

Responsibilities

Own and enhance the Kubernetes scheduler for accelerator fleets
Scale the control plane to support expansive clusters, identifying and resolving bottlenecks
Design and build essential cluster services relied upon by all workloads
Develop and manage custom controllers and operators
Collaborate with teams to translate workload needs into platform functionalities
Engage with cloud providers to address feature requests and escalate issues
Lead incident responses and establish improvement processes to minimize failure recurrence

Benefits

Hybrid work model with a minimum in-office requirement
Visa sponsorship opportunities
Inclusive application encouragement for underrepresented groups
Focus on ethical and socially impactful AI development
Supportive recruiting to combat scams and ensure candidate safety

Full Job Description

About the role

Anthropic runs some of the largest Kubernetes clusters in the industry. We have fleets of hundreds of thousands of nodes across multiple cloud providers and datacenters to train, research, and serve frontier AI models. The Kubernetes Platform team owns the Kubernetes control plane that makes those clusters work.

We are operating at a scale where the defaults stop working. We own the scheduler and extend it to place topology-sensitive ML workloads across thousands of accelerators at once. We scale the control plane itself - apiserver, etcd, controllers - so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure.

We make sure the control plane is fast, correct, and always available. Your work will directly determine whether Anthropic can keep reliably and safely training frontier models as our compute footprint continues to grow.
Key responsibilities

Own, operate, and extend the Kubernetes scheduler for Anthropic's accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption
Scale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds us
Design, build, and operate core cluster services such as service discovery that every workload in the fleet depends on
Build and maintain custom controllers, operators, and CRDs
Partner with research, training, and inference to understand workload shapes and turn their requirements into platform capabilities
Collaborate with cloud providers on required features and escalations
Participate in on-call, lead incident response, and design processes (postmortems, runbooks, SLOs) that help the team avoid repeating failures

Minimum qualifications

Significant software engineering experience building and operating production distributed systems
Proficiency in at least one systems-appropriate language (e.g., Go, Python, Rust, or C++)
Deep, hands-on Kubernetes experience (well beyond "user of") into scheduler, controllers, apiserver, or operating large multi-tenant clusters
Demonstrated ability to debug complex issues across the stack, from API behavior down to node and network-level root causes
A track record of designing for reliability, correctness, and clear failure semantics in systems other engineers depend on
Strong written and verbal communication; comfort building consensus with internal stakeholders

Preferred qualifications

Experience with Kubernetes internals or contributions: kube-scheduler / scheduling framework, apiserver, etcd, client-go, controller-runtime, or similar
Experience building or operating cluster schedulers or batch systems (e.g., Kueue, Volcano, Slurm, or in-house equivalents)
Background scaling control planes or coordination systems (etcd, ZooKeeper, Consul, or large DNS/service-mesh deployments)
Familiarity with ML infrastructure: GPUs, TPUs, or Trainium; gang scheduling; topology-aware placement; collective networking such as NCCL
Experience with GCP and/or AWS, including GKE/EKS internals and Infrastructure as Code
Low-level systems experience such as Linux kernel tuning, cgroups, or eBPF
8+ years of relevant industry experience, including time leading large, ambiguous infrastructure projects

The annual compensation range for this role is listed below.

For sales roles, the range provided is the role's On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.

Annual Salary:

$320,000-$405,000 USD

Logistics

Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience

Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience

Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position

Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.

We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.

Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from [redacted].com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links-visit anthropic.com/careers directly for confirmed position openings.

About Anthropic

Anthropic is an artificial intelligence research lab that focuses on developing AI systems that are safe, reliable, and trustworthy. The company was founded in 2019 by Dr. Yoshua Bengio, a leading AI researcher and winner of the Turing Award. Anthropic's research is focused on developing AI systems that can learn from small amounts of data, reason about complex systems, and interact with humans in a natural way. The company is based in New York City and has a team of experienced AI researchers and engineers.

Learn more about Anthropic

Size

50 employees

Industry

Information Technology

Founded

2019

* Ladders Estimates

Similar Jobs

Staff Software Engineer, Capacity Engineering
$177K — $364K *
Pinterest
Remote
Today
Staff Software Engineer, Capacity Engineering
$177K — $364K *
Pinterest
San Francisco, CA 94112 (San Francisco County)
Today
Senior Staff Machine Learning Engineer, Consumer
$242K — $357K *
DoorDash
San Francisco, CA 94112 (San Francisco County)
Today
Staff Software Engineer- Growth Performance Marketing
$177K — $364K *
Pinterest
Remote
Today
Senior Staff Software Engineer, Serverless
$262K — $365K *
Google
Sunnyvale, CA 94087 (Santa Clara County)
Today
Staff Software Engineer, Mapping
$185K — $335K *
General Motors
Remote
Reposted 2 days ago

Get Ready For Your
Next Interview

More Jobs at Anthropic

Manager, Account Executive - Enterprise Sales
$360K — $500K+*
San Francisco, CA 94112 (San Francisco County)
Yesterday
Enterprise Technology
In-Person
Manager, Account Executive - Enterprise Sales
$360K — $500K+*
New York, NY 10025 (New York County)
Yesterday
Enterprise Technology
In-Person
Data Scientist, GTM
$275K — $370K *
San Francisco, CA 94112 (San Francisco County)
Yesterday
Business Services
In-Person
Data Scientist, GTM
$275K — $370K *
New York, NY 10025 (New York County)
Yesterday
Business Services
In-Person
Data Scientist, Marketing
$275K — $370K *
San Francisco, CA 94112 (San Francisco County)
Reposted 2 days ago
Consumer Technology
In-Person

More Information Technology Jobs

Business Development Director
$300K — $345K + $120K bonus *
Tier1 IT Services Firm
Kansas City, MO 64116 (Clay County)
6 days ago
Client Partner / Business Developemnt - Banking
$250K — $320K + $70K bonus *
IT Services Firm (client of TechLink Systems)
New York, NY 10001 (New York County)
6 days ago
Customer Support
Confidential Company
Austin, TX 78701 (Travis County)
2 weeks ago
Sr Assoc, Cyber Sec ThreatMgmt - Detection Engineer
$88K — $151K *
Northern Trust
Naperville, IL 60540 (Dupage County)
Today
Global Director – Vulnerability Management & Security Configuration
$164K — $288K *
Northern Trust
Chicago, IL 60629 (Cook County)
Today

Find similar Staff Software Engineer, Kubernetes Platform jobs:

Nationwide San Francisco, CA

Staff Software Engineer, Kubernetes Platform

Job Overview by Ladders

Full Job Description

Get Ready For Your Next Interview

Find similar Staff Software Engineer, Kubernetes Platform jobs:

Get Ready For Your
Next Interview