Software Engineer, Compute Foundations
OpenAI- Compensation
- $255k–$490k Published range · Top quartile for Engineering (773 listings)
- Location
- Remote - US Remote eligibility
- Employment
- Full-time Mid-level
Role skills
Apply on jobs.ashbyhq.com
About the job
About the Team
Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters.
About the Role
You will build distributed systems that provision, configure, and manage compute throughout its lifecycle, connecting global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong.
In this role, you will
- Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows.
- Define APIs and resource models for lifecycle operations across hardware platforms and providers.
- Build provisioning and configuration services coordinating network boot, hardware management interfaces, and deployment of firmware, OS images, drivers, and host configuration.
- Develop lifecycle management for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning.
- Design reliable reconciliation and recovery with staged rollouts that limit disruption.
- Improve control-plane throughput, API latency, and time-to-desired-state while respecting provider limits.
- Integrate new sites and GPU hardware generations into the platform.
You might thrive in this role if you
- Have strong software engineering fundamentals and experience owning production distributed systems or infrastructure services.
- Have experience with Kubernetes APIs and reconciliation.
- Understand bare-metal node lifecycle (PXE, DHCP/DNS, BMCs, firmware, Linux, drivers, images, configuration management).
- Can design reliable APIs and asynchronous workflows, reasoning about concurrency, consistency, idempotency, and failures.
- Can diagnose reliability and performance problems across service, OS, and machine boundaries.
- Work effectively across engineering specialties and communicate technical tradeoffs clearly.
Bonus points if you
- Have built infrastructure control planes coordinating operations across multiple sites or regions.
- Have worked with GPU or HPC infrastructure.
- Have integrated multiple hardware platforms or providers into a common service model.
Compensation
Salary range: $255K – $490K. Offers equity.
Skills & tags
What you can verify before applying
Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.