Staff Infrastructure Engineer, Cluster Infrastructure
Anthropic- Compensation
- £325k–£485k Published range
- Location
- Hybrid - London, UK, at least 25% in office Remote eligibility
- Employment
- Full-time Staff / Principal
About the job
Anthropic is an AI safety company building reliable, interpretable, and steerable AI systems. The Infrastructure organization is foundational to this mission, and the Cluster Infra team owns the full lifecycle of compute clusters, building agent-driven automation for provisioning and lifecycle management across cloud providers and datacenters.
Responsibilities
- Own the technical strategy and roadmap for agent-driven cluster lifecycle management (provisioning, updates, decommissioning).
- Partner across teams to ensure new compute capacity is ingested on time.
- Align with partner teams on physical build-out and leverage cloud solutions for high-bandwidth inter-cluster connectivity.
- Collaborate with security owners to ensure clusters are provisioned secure-by-default.
- Define and drive strategy on cluster scalability, homogeneity, and fault tolerance.
- Work with cloud providers and internal teams to shape long-term compute, data, and infrastructure strategy.
- Establish operational-excellence practices: incident response, postmortem culture, on-call health.
- Support engineer growth through technical mentorship and coaching.
Qualifications
Minimum
- Deep expertise in distributed systems, reliability, and cloud platforms (e.g., Kubernetes, IaC, AWS/GCP/Azure).
- Strong proficiency in at least one systems language (e.g., Rust, Go, or Python) and IaC proficiency with Terraform.
- Track record of leading complex, multi-quarter technical initiatives spanning multiple teams or systems.
- Ability to build alignment across senior stakeholders and communicate effectively.
Preferred
- 10+ years of software engineering experience, including time as a technical lead.
- Experience operating large-scale compute infrastructure at hyperscale (100+ clusters, 10K+ nodes).
- Depth in Kubernetes internals, cluster provisioning/management systems, or cluster orchestration systems.
- Experience with cloud networking (VPC design, peering, Transit Gateway, Cloud Interconnect, etc.).
- Experience with cluster and host networking (CNI, eBPF, service mesh, mTLS).
- Experience with cluster security (pod security standards, RBAC, IAM, hardening).
- Deep experience with infrastructure-as-code (Terraform, Atlantis) and workflow orchestration (Temporal, Argo Workflows).
Compensation
Annual Salary: £325,000—£485,000 GBP.
Benefits
Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a collaborative office space.
Application Instructions
Apply via the Greenhouse link. Anthropic sponsors visas and will make every reasonable effort to get you a visa if an offer is made. The role is subject to a hybrid policy requiring at least 25% time in an office.
Skills & tags
Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.