Skip to main content

Engineering Manager, Kubernetes Infrastructure (Bare Metal)

CoreWeave
Hybrid - Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WAUpdated 13h ago
Base salary
$182k–$242k
Published base salary range
Location
Hybrid - Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA
Remote eligibility
Employment
Full-time
Lead / Manager
Role family
Engineering
AI / ML
Apply on coreweave.com
Job actionsApply now
Job actionsApply now

About the job

About the role

CoreWeave is looking for an Engineering Manager to lead a team building and operating Kubernetes infrastructure on bare metal. This team sits close to the core of the platform and is responsible for the reliability, scalability, and operational excellence of the systems that power high-performance AI and ML workloads.

What you'll do

  • Lead a team of engineers responsible for Kubernetes infrastructure running on bare metal.
  • Set clear goals, priorities, and execution plans for the team.
  • Partner with senior ICs and adjacent teams on the roadmap for cluster lifecycle management, upgrades, reliability, observability, and infrastructure automation.
  • Improve operational excellence across incident response, on-call health, root-cause analysis, and service ownership.
  • Drive engineering best practices for safe change management, testing, rollout quality, and production readiness.
  • Support the design and operation of platform capabilities for provisioning, patching, upgrades, scaling, and troubleshooting of Kubernetes clusters.
  • Build strong cross-functional relationships with compute, networking, storage, security, and product stakeholders.
  • Hire, coach, and develop engineers while creating a high-accountability, high-trust team culture.

Who you are

  • Experience managing an infrastructure, platform, or SRE-oriented engineering team.
  • Strong technical depth in Kubernetes, distributed systems, and production infrastructure.
  • Experience operating Kubernetes in complex environments, ideally including bare metal, hybrid, or highly performance-sensitive systems.
  • Familiarity with cluster lifecycle management, including provisioning, upgrades, node operations, observability, and reliability engineering.
  • Track record of improving team execution, engineering quality, and operational maturity.
  • Experience leading incident response cultures and driving follow-through on reliability improvements.
  • Strong partnership skills across engineering, product, and operations functions.
  • Ability to coach engineers at different levels and create clarity in ambiguous or fast-scaling environments.
  • Strong written and verbal communication.

Compensation

The base salary range for this role is $182,000 to $242,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. In addition to base salary, the total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).

Benefits

Medical, dental, and vision insurance - 100% paid by CoreWeave; company-paid life insurance; short and long-term disability insurance; flexible spending account; health savings account; tuition reimbursement; employee stock purchase program; mental wellness benefits through Spring Health; family-forming support provided by Carrot; paid parental leave; flexible, full-service childcare support with Kinside; 401(k) with a generous employer match; flexible PTO; catered lunch each day in our office and data center locations; a casual work environment.

Application

Apply via the provided careers link. For reasonable accommodation, contact careers@coreweave.com.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.