Skip to main content

Staff+ Software Engineer, Kubernetes Platform

Anthropic
Hybrid - San Francisco, New York City, or Seattle, at least 25% in officeUpdated 1d ago
Compensation
$320k–$485k
Published range · Top quartile for Engineering (582 listings)
Location
Hybrid - San Francisco, New York City, or Seattle, at least 25% in office
Remote eligibility
Employment
Full-time
Staff / Principal
Role family
Engineering
AI / ML
Apply on job-boards.greenhouse.io
Job actionsApply now
Job actionsApply now

About the job

About the role

Anthropic runs one of the industry's largest AI compute fleets, spanning multiple cloud providers and datacenters, to train, research, and serve frontier AI models. The Kubernetes Platform team owns the control plane that makes these fleets work, operating at a scale where defaults stop working. The team owns the scheduler, scales the control plane (apiserver, etcd, controllers), and builds core cluster services like service discovery.

Key responsibilities

  • Own, operate, and extend the Kubernetes scheduler for accelerator fleets, including custom scheduling plugins for gang scheduling, topology awareness, and preemption.
  • Scale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits.
  • Design, build, and operate core cluster services such as service discovery.
  • Build and maintain custom controllers, operators, and CRDs.
  • Partner with research, training, and inference teams to turn workload requirements into platform capabilities.
  • Collaborate with cloud providers on required features and escalations.
  • Participate in on-call, lead incident response, and design processes (postmortems, runbooks, SLOs).

Minimum qualifications

  • Significant software engineering experience building and operating production distributed systems.
  • Proficiency in at least one systems-appropriate language (e.g., Go, Python, Rust, or C++).
  • Deep, hands-on Kubernetes experience (beyond 'user of') into scheduler, controllers, apiserver, or operating large multi-tenant clusters.
  • Demonstrated ability to debug complex issues across the stack.
  • Track record of designing for reliability, correctness, and clear failure semantics.
  • Strong written and verbal communication.

Preferred qualifications

  • Experience with Kubernetes internals or contributions (kube-scheduler, apiserver, etcd, client-go, controller-runtime).
  • Experience building or operating cluster schedulers or batch systems (Kueue, Volcano, Slurm).
  • Background scaling control planes or coordination systems (etcd, ZooKeeper, Consul).
  • Familiarity with ML infrastructure: GPUs, TPUs, Trainium; gang scheduling; topology-aware placement; collective networking such as NCCL.
  • Experience with GCP and/or AWS, including GKE/EKS internals and Infrastructure as Code.
  • Low-level systems experience such as Linux kernel tuning, cgroups, or eBPF.
  • 10+ years of relevant industry experience.

Compensation

Annual Salary: $320,000—$485,000 USD.

Logistics

Minimum education: Bachelor's degree or equivalent. Location-based hybrid policy: all staff in one of our offices at least 25% of the time. Visa sponsorship: We do sponsor visas.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.