Skip to main content

Applied AI Engineer, Inference

CoreWeave
Hybrid - Bellevue, WA; San Francisco, CA; Sunnyvale, CAUpdated 14h ago
Base salary
$188k–$275k
Published base salary range
Location
Hybrid - Bellevue, WA; San Francisco, CA; Sunnyvale, CA
Remote eligibility
Employment
Full-time
Mid-level
Role family
Engineering
AI / ML
Apply on coreweave.com
Job actionsApply now
Job actionsApply now

About the job

About the Role

CoreWeave is The Essential Cloud for AI, a publicly traded company (Nasdaq: CRWV) delivering a platform for building and scaling AI. The Inference team is responsible for high-performance model serving. This Applied AI Engineer role focuses on understanding, measuring, and improving real-world performance of the inference platform, with initial responsibilities in benchmarking, optimization, and workload-driven research.

Responsibilities

  • Build and maintain benchmarking workflows measuring latency, throughput, quality regressions, and cost.
  • Benchmark the inference stack against realistic customer workloads and external provider baselines.
  • Profile model-serving behavior across frameworks, runtimes, and hardware to find bottlenecks in prefill, decode, KV cache usage, batching, graph capture, quantization, and related systems.
  • Drive targeted optimizations for customer and product workloads, tuning serving configurations and validating changes.
  • Design and run experiments on model-serving techniques such as quantization, speculative decoding, caching, and routing.
  • Partner with inference platform engineers to productionize improvements and establish repeatable performance testing workflows.
  • Produce technical writeups and recommendations on model configurations, runtime choices, hardware allocation, and deployment strategies.

Qualifications

  • 4+ years of experience in machine learning, systems, performance engineering, or adjacent applied engineering work.
  • Strong programming skills in Python and comfort in production engineering environments.
  • Experience running empirical evaluations, benchmarks, or experiments and translating results into engineering decisions.
  • Familiarity with LLM inference systems and tools such as vLLM, SGLang, TensorRT-LLM, or similar model-serving stacks.
  • Understanding of tradeoffs in latency, throughput, batching, GPU utilization, quantization, and quality regression analysis.
  • Ability to work across model, systems, and product boundaries.
  • Strong written communication and a bias toward making technical work legible and reproducible.

Compensation & Benefits

The base salary range for this role is $188,000 to $275,000, determined by job-related knowledge, skills, experience, and market location. Total rewards include a discretionary bonus, equity awards, and a comprehensive benefits program. Benefits include medical, dental, and vision insurance (100% paid by CoreWeave), company-paid life insurance, short- and long-term disability insurance, tuition reimbursement, mental wellness benefits, family-forming support, paid parental leave, childcare support, 401(k) with employer match, flexible PTO, and catered lunch in office and data center locations.

Application Instructions

This position requires access to export controlled information; applicants must be a U.S. person, eligible to access without authorization, or eligible and reasonably likely to obtain the required export authorization.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.