Skip to main content

Software Engineer, Compute Foundations

OpenAI
Remote - USUpdated 5d ago
Compensation
$255k–$490k
Published range · Top quartile for Engineering (773 listings)
Location
Remote - US
Remote eligibility
Employment
Full-time
Mid-level
Role family
Engineering
AI / ML
Apply on jobs.ashbyhq.com
Job actionsApply now
Job actionsApply now

About the job

About the Team

Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters.

About the Role

You will build distributed systems that provision, configure, and manage compute throughout its lifecycle, connecting global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong.

In this role, you will

  • Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows.
  • Define APIs and resource models for lifecycle operations across hardware platforms and providers.
  • Build provisioning and configuration services coordinating network boot, hardware management interfaces, and deployment of firmware, OS images, drivers, and host configuration.
  • Develop lifecycle management for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning.
  • Design reliable reconciliation and recovery with staged rollouts that limit disruption.
  • Improve control-plane throughput, API latency, and time-to-desired-state while respecting provider limits.
  • Integrate new sites and GPU hardware generations into the platform.

You might thrive in this role if you

  • Have strong software engineering fundamentals and experience owning production distributed systems or infrastructure services.
  • Have experience with Kubernetes APIs and reconciliation.
  • Understand bare-metal node lifecycle (PXE, DHCP/DNS, BMCs, firmware, Linux, drivers, images, configuration management).
  • Can design reliable APIs and asynchronous workflows, reasoning about concurrency, consistency, idempotency, and failures.
  • Can diagnose reliability and performance problems across service, OS, and machine boundaries.
  • Work effectively across engineering specialties and communicate technical tradeoffs clearly.

Bonus points if you

  • Have built infrastructure control planes coordinating operations across multiple sites or regions.
  • Have worked with GPU or HPC infrastructure.
  • Have integrated multiple hardware platforms or providers into a common service model.

Compensation

Salary range: $255K – $490K. Offers equity.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.