Skip to main content

Senior Software Engineer, AI Runtime

Databricks
Hybrid - Mountain View or San FranciscoUpdated 5d ago
Total compensation
$160k–$225k
Published total compensation range
Location
Hybrid - Mountain View or San Francisco
Remote eligibility
Employment
Full-time
Senior
Role family
Engineering
AI / ML
Apply on databricks.com
Job actionsApply now
Job actionsApply now

About the job

About the role

Databricks is the Data and AI company, building the world's best data and AI infrastructure platform. AI Runtime (AIR) is our managed platform for large-scale GPU training and fine-tuning, providing on-demand access to fleets of accelerators and a serverless experience for multi-node jobs. As a Senior Software Engineer for AI Runtime, you will build and scale the systems that make large-scale training fast, reliable, and effortless.

What you'll do

  • Drive the architecture and evolution of AIR's managed GPU training platform, delivering scalable, high-throughput, and resilient training across fleets spanning thousands of accelerators.
  • Solve the hardest problems in large-scale training, including multi-node orchestration, distributed parallelism strategies, GPU scheduling, high-throughput data loading, and checkpoint/restore for long-running jobs.
  • Push GPU efficiency and training performance, raising utilization and lowering cost per training run.
  • Build resilience and observability foundations to keep multi-node jobs healthy.
  • Partner with product, research, and platform teams to shape APIs, CLI, and developer experience.
  • Lead end-to-end engineering efforts from design through production rollout.
  • Mentor other engineers and contribute to Databricks' technical direction in AI training infrastructure.

What we look for

  • 5+ years of experience building and operating large-scale distributed systems, with experience in GPU training infrastructure, high-performance computing, or ML systems.
  • Experience with distributed training frameworks (such as PyTorch, FSDP, DeepSpeed, or Megatron) and parallelism strategies.
  • Strong understanding of training resilience patterns, including checkpointing, failure detection, and automatic recovery.
  • Solid grasp of GPU performance fundamentals, including accelerator architecture, high-speed interconnects (NVLink, InfiniBand/RoCE), and collective communication.
  • Experience building and operating managed, multi-tenant platform products in the cloud with clear SLAs/SLOs.
  • Strong foundation in algorithms, data structures, and system design.
  • BS in Computer Science or related field (MS or PhD preferred).

Compensation

Local Pay Range: $160,000—$225,000 USD. Total compensation may include annual performance bonus, equity, and benefits.

Benefits

Databricks offers comprehensive benefits and perks. Specific details vary by region.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.