Skip to main content

Performance Engineer, Inference Systems

Anthropic
Hybrid - San Francisco, CA, New York City, NY, or Seattle, WA, 25% in officeUpdated 1d ago
Compensation
$350k–$850k
Published range · Top quartile for Engineering (545 listings)
Location
Hybrid - San Francisco, CA, New York City, NY, or Seattle, WA, 25% in office
Remote eligibility
Employment
Full-time
Mid-level
Role family
Engineering
AI / ML
Role skills
Apply on job-boards.greenhouse.io
Job actionsApply now
Job actionsApply now

About the job

About Anthropic

Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We are a growing team of researchers, engineers, policy experts, and business leaders building beneficial AI systems.

About the Role

Anthropic's inference fleet serves Claude to millions of users. The Inference System Dynamics team is responsible for understanding the whole system and holding it to a high bar across throughput, latency, reliability, and correctness. We measure fleet performance against theoretical frontiers, run cross-layer investigations, and own correctness checks.

Key Responsibilities

  • Run cross-layer performance investigations across throughput, latency, and reliability, sizing gaps and identifying root causes.
  • Own and improve the correctness evaluation pipeline that validates model output quality across hardware platforms, numerics, and serving configurations.
  • Build observability, dashboards, and modeling tools to make performance metrics legible across the stack.
  • Partner with kernel, serving, routing, autoscaling, and capacity teams to prioritize and land optimizations.
  • Stack-rank opportunities by impact and effort.

Minimum Qualifications

  • Hands-on performance engineering experience: profiling, roofline analysis, latency/throughput optimization, and root-cause investigation in complex production systems.
  • Proficiency in Python, with ability to read, instrument, and contribute to large production codebases.
  • Solid data analysis skills (e.g., SQL, pandas) to turn raw telemetry into findings.
  • Ability to communicate quantitative results clearly in writing.
  • Genuine interest in correctness as an engineering discipline.

Preferred Qualifications

  • Experience with ML systems, especially training or inference infrastructure or LLM serving stacks.
  • Familiarity with GPU/TPU/accelerator performance concepts.
  • Experience with reliability engineering for high-throughput services.
  • Experience with model evaluation or numerical regression-detection pipelines.
  • Experience building observability or telemetry for distributed systems.

Compensation

Annual Salary: $350,000—$850,000 USD.

Location and Hybrid Policy

Location: San Francisco, CA | New York City, NY | Seattle, WA. Hybrid: we expect all staff to be in one of our offices at least 25% of the time.

Visa Sponsorship

We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.