Skip to main content

Staff Software Engineer, Inference

CoreWeave
Hybrid - Sunnyvale, CA / Bellevue, WAUpdated 15h ago
Base salary
$188k–$275k
Published base salary range
Location
Hybrid - Sunnyvale, CA / Bellevue, WA
Remote eligibility
Employment
Full-time
Staff / Principal
Role family
Engineering
AI / ML
Apply on coreweave.com
Job actionsApply now
Job actionsApply now

About the job

CoreWeave is The Essential Cloud for AI™, a specialized cloud provider delivering GPU compute for AI workloads. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025.

The Inference team builds and operates CoreWeave's Kubernetes-native inference platform, powering low-latency, high-throughput AI workloads at massive scale. The team is responsible for request routing, scheduling, GPU resource management, and system-wide optimizations.

As a Staff Software Engineer (IC5) on the Inference team, you will act as a technical leader driving architecture, performance, and reliability across multiple services and teams. You will lead cross-team design initiatives, optimize inference performance (latency, throughput, GPU utilization), and improve system reliability at scale. You will work deeply in distributed systems and Kubernetes-based infrastructure, focusing on scheduling, batching, and memory optimization.

Qualifications

  • 8–12+ years of experience building and operating large-scale distributed systems or cloud platforms
  • Proven experience leading cross-team technical initiatives
  • Strong programming skills in Go, Python, or C++
  • Deep expertise in Kubernetes at production scale
  • Strong understanding of distributed systems, networking, and performance optimization
  • Experience designing and operating low-latency, high-throughput systems with strict P95/P99 latency requirements
  • Hands-on experience with inference systems, including batching, caching, and memory optimization
  • Experience improving system performance using metrics-driven approaches
  • Familiarity with mixed precision (BF16, FP8) and streaming inference workloads

Preferred: Experience with inference frameworks such as vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe; GPU systems and performance optimization (CUDA, NCCL, RDMA, NUMA); leading multi-team initiatives; exposure to large-scale AI/ML infrastructure.

Compensation

The base salary range for this role is $188,000 to $275,000. Total rewards include a discretionary bonus, equity awards, and a comprehensive benefits program.

Benefits

  • Medical, dental, and vision insurance (100% paid by CoreWeave)
  • Company-paid life insurance
  • Short and long-term disability insurance
  • Tuition reimbursement
  • Employee Stock Purchase Program (ESPP)
  • Mental wellness benefits
  • Family-forming support
  • Paid parental leave
  • Childcare support
  • 401(k) with generous employer match
  • Flexible PTO
  • Catered lunch in office and data center locations

Application

Apply via the provided link. For accommodations, contact careers@coreweave.com.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.