Skip to main content

Staff + Senior Software Engineer, Inference Infrastructure

Anthropic
Hybrid - San Francisco, CA | New York City, NY | Seattle, WA, at least 25% in officeUpdated 34d ago
Compensation
$320k–$485k
Published range · Top quartile for Engineering (813 listings)
Location
Hybrid - San Francisco, CA | New York City, NY | Seattle, WA, at least 25% in office
Remote eligibility
Employment
Full-time
Staff / Principal
Role family
Engineering
AI / ML
Apply on job-boards.greenhouse.io
Job actionsApply now
Job actionsApply now

About the job

Anthropic is a public benefit corporation headquartered in San Francisco, focused on creating reliable, interpretable, and steerable AI systems. The Inference team builds and maintains the critical systems that serve Claude to millions of users worldwide, operating the industry's largest compute-agnostic inference deployments. The team handles the entire stack from intelligent request routing to fleet-wide orchestration across diverse AI accelerators, with a dual mandate: maximizing compute efficiency to serve customer growth and enabling breakthrough research by providing high-performance inference infrastructure. Key responsibilities include designing, building, and maintaining distributed systems; developing resilient systems that adapt in real time; developing intelligent request routing, load balancing, and traffic management across thousands of accelerators; maximizing compute efficiency through autoscaling and orchestration; building production-grade deployment pipelines; providing high-performance inference infrastructure for researchers; and integrating new AI accelerator platforms. Minimum qualifications include significant software engineering experience with distributed systems, results-oriented mindset, willingness to pick up slack, desire to learn about machine learning systems and infrastructure, thriving in environments where technical excellence drives business results, and caring about societal impacts. Preferred qualifications include experience with high-performance large-scale distributed systems, implementing and deploying machine learning systems at scale, load balancing/request routing/traffic management, LLM inference optimization, Kubernetes and cloud infrastructure (AWS, GCP, Azure), and proficiency in Python or Rust. Representative projects include designing intelligent routing algorithms, autoscaling compute fleet, building deployment pipelines, contributing to new inference features, supporting new model architectures, analyzing observability data, and managing multi-region deployments. The annual salary range is $320,000—$485,000 USD. Minimum education is a Bachelor's degree or equivalent. The role is hybrid, expecting staff in offices at least 25% of the time. Applications are reviewed on a rolling basis with no deadline.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.