Staff+ Software Engineer, Capacity Engineering
Anthropic- Compensation
- $320k–$485k Published range · Top quartile for Engineering (505 listings)
- Location
- Hybrid - San Francisco, New York City, or Seattle, 25%+ in office Remote eligibility
- Employment
- Full-time Staff / Principal
About the job
Anthropic is a public benefit corporation building reliable, interpretable, and steerable AI systems. The Capacity Engineering team owns the data, tooling, and operational systems that let Anthropic plan, measure, and maximize utilization across first-party and third-party compute. As a Staff+ Software Engineer on this team, you will build production systems that power this work: data pipelines that ingest and normalize telemetry from heterogeneous cloud environments, observability tooling for real-time fleet health, and performance instrumentation that measures how efficiently every major workload uses its hardware.
This is a pipeline role feeding four areas: data platform (pipelines ingesting occupancy and utilization telemetry from Kubernetes clusters, normalizing billing and usage across cloud providers, serving BigQuery tables), planning (cluster health tooling, capacity planning platforms, alerting on occupancy drops), efficiency (instrumenting utilization across training, inference, and eval systems, building benchmarking infrastructure, establishing per-config baselines), and attribution/forecasting (reconciling CSP billing exports, attributing spend to workloads, turning demand signals into a defensible compute plan).
Key responsibilities
- Build the planning and allocation stack — tools for capacity allocation, cross-region/cross-provider placement, guardrails, queueing, and occupancy KPIs.
- Drive efficiency programs: stranding and rightsizing, unused capacity recovery, job-level utilization across training, inference, and eval.
- Own attribution and forecasting — reconcile billing across ten-plus providers, attribute spend, and produce a defensible compute plan.
- Build the data platform: pipelines ingesting occupancy, utilization, and cost into BigQuery, with ownership of completeness, latency SLOs, and gap detection.
- Operate Kubernetes-native systems at scale — collection agents, workload labeling, taint/reservation/scheduling behavior.
- Treat the output as a product, not a pipeline — gather requirements, define schema contracts, design for consumers from research engineers to a CFO, including on-call and SLOs.
What you bring
- Strong track record building and operating production systems; hands-on engineering with a devops flavor.
- Python and SQL at production quality (most pipeline code is Python; presentation layer is BigQuery SQL, including table-valued functions and views).
- Deep experience with at least one major cloud provider (AWS, GCP, or Azure) and its operations.
- Experience with observability tooling stack: Prometheus, PromQL, and Grafana, including writing recording rules and building monitoring.
- Ability to gather your own requirements and work across organizational boundaries in an ambiguous environment.
Preferred qualifications
- Experience with capacity planning, resource management, or cost attribution systems at a hyperscaler or large-scale ML environment.
- Scheduling and packing efficiency experience, or profiling-driven optimization of large distributed workloads.
- Multi-cloud data ingestion experience, especially normalizing billing exports, reservation APIs, on-demand capacity reservations, commitments, and vendor telemetry.
- Total cost of ownership and forecasting experience, including decomposing whether infrastructure growth is causal or correlated with business drivers.
- Accelerator infrastructure familiarity (GPU metrics like DCGM, TPU utilization, Trainium power/utilization metrics, or ML training/inference systems at the hardware level).
- Experience building internal data products with self-service access, schema contracts, API serving, documentation, and discoverability.
- Storage efficiency, retention, and lifecycle program experience at exabyte scale.
Compensation: Annual Salary: $320,000—$485,000 USD.
Logistics: Minimum education: Bachelor’s degree or equivalent combination of education, training, and/or experience. Required field of study: relevant to the role. Years of experience required will correlate with internal job level requirements. Location-based hybrid policy: all staff expected to be in one of our offices at least 25% of the time. Visa sponsorship: We do sponsor visas! We will make every reasonable effort to get you a visa if we make an offer.
Skills & tags
Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.