Senior Software Engineer, AI Runtime
Databricks- Total compensation
- $160k–$225k Published total compensation range
- Location
- Hybrid - Mountain View or San Francisco Remote eligibility
- Employment
- Full-time Senior
About the job
About the role
Databricks is the Data and AI company, building the world's best data and AI infrastructure platform. AI Runtime (AIR) is our managed platform for large-scale GPU training and fine-tuning, providing on-demand access to fleets of accelerators and a serverless experience for multi-node jobs. As a Senior Software Engineer for AI Runtime, you will build and scale the systems that make large-scale training fast, reliable, and effortless.
What you'll do
- Drive the architecture and evolution of AIR's managed GPU training platform, delivering scalable, high-throughput, and resilient training across fleets spanning thousands of accelerators.
- Solve the hardest problems in large-scale training, including multi-node orchestration, distributed parallelism strategies, GPU scheduling, high-throughput data loading, and checkpoint/restore for long-running jobs.
- Push GPU efficiency and training performance, raising utilization and lowering cost per training run.
- Build resilience and observability foundations to keep multi-node jobs healthy.
- Partner with product, research, and platform teams to shape APIs, CLI, and developer experience.
- Lead end-to-end engineering efforts from design through production rollout.
- Mentor other engineers and contribute to Databricks' technical direction in AI training infrastructure.
What we look for
- 5+ years of experience building and operating large-scale distributed systems, with experience in GPU training infrastructure, high-performance computing, or ML systems.
- Experience with distributed training frameworks (such as PyTorch, FSDP, DeepSpeed, or Megatron) and parallelism strategies.
- Strong understanding of training resilience patterns, including checkpointing, failure detection, and automatic recovery.
- Solid grasp of GPU performance fundamentals, including accelerator architecture, high-speed interconnects (NVLink, InfiniBand/RoCE), and collective communication.
- Experience building and operating managed, multi-tenant platform products in the cloud with clear SLAs/SLOs.
- Strong foundation in algorithms, data structures, and system design.
- BS in Computer Science or related field (MS or PhD preferred).
Compensation
Local Pay Range: $160,000—$225,000 USD. Total compensation may include annual performance bonus, equity, and benefits.
Benefits
Databricks offers comprehensive benefits and perks. Specific details vary by region.
Skills & tags
Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.