Skip to main content

Lead Machine Learning Engineer - ML Infrastructure

Samsara
Remote - USUpdated 15d ago
Base salary
$196k–$270k
Published base salary range
Location
Remote - US
Remote eligibility
Employment
Full-time
Lead / Manager
Role family
Engineering
B2B SaaS
Apply on samsara.com
Job actionsApply now
Job actionsApply now

About the job

About the role

Samsara is the industry leader in AI for physical operations. We're hiring a Lead Machine Learning Infrastructure Engineer to serve as the technical anchor for ML infrastructure across Samsara's Safety AI organization. You will own the architecture and evolution of our end-to-end ML platform — spanning training, experimentation, inference, and edge deployment across more than 2M deployed devices — and be the connective tissue between applied ML teams, security, and data platform. This role is not an execution role sitting under a technical lead; you are the technical lead. Your decisions shape platform direction, unblock multiple product teams, and translate directly into real-world safety outcomes for the industries that run our world. We are open to calibrating this role at Staff, Senior Staff, or Principal level depending on the candidate's scope of experience — what matters most is end-to-end platform ownership and the ability to operate as the technical anchor for the organization.

In this role, you will

  • Set the technical strategy and own end-to-end delivery of Samsara's ML platform (training, experimentation, batch/online inference, edge) — making architectural decisions and being the accountability point across all platform layers for multiple Safety AI product teams.
  • Drive the design, launch, and iteration of Safety AI features (CV models, EcoDriving insights, LLM-based reporting) — not just enabling others to ship, but co-owning outcomes including safety metrics, reliability, and cost at production scale.
  • Design and operate scalable online and batch inference systems (Ray, Spark), including deployment patterns, observability, SLOs, and unified training-to-production workflows. Partner with firmware and edge teams to package, validate, and deploy models to Samsara devices, and build feedback loops from edge to cloud for continuous improvement.
  • Own reliability, observability, and security for ML systems across cloud and edge, including on-call practices, incident response, and infrastructure hardening. Own or co-own end-to-end technical delivery for high-priority or high-risk initiatives, from modeling and system design through production rollout.
  • Be the technical authority for ML infrastructure architecture across Safety AI — setting direction that cross-functional teams (applied ML, firmware, security, data platform) execute against, mentoring senior engineers and applied scientists, and ensuring platform decisions are made at the right level of abstraction with the right trade-offs.
  • Drive strong developer experience through documentation and best practices, while contributing to and representing Samsara in open source communities (Ray, Spark, RayDP).

Minimum requirements

  • 10+ years in machine learning engineering, with demonstrated tech lead ownership of at least two major ML platform domains (distributed training, data/research infrastructure, cloud inference, or feature engineering) serving multiple product teams at scale.
  • Proven record of shipping ML-powered features end-to-end — from design through production and iteration — with measurable impact on product or business metrics (not just building internal tooling).
  • Hands-on Ray and Kubernetes expertise in production environments; Spark experience strongly preferred.
  • Able to be a credible peer to the most senior engineers on the team.
  • Deep understanding of ML fundamentals beyond pipelines: evaluation methodology, dataset design, ablation, drift, and the ability to review and redirect modeling approaches — you bridge research and engineering, not just serve them.
  • Demonstrated cross-org technical leadership around platform decisions, and influencing roadmap and go/no-go calls based on throughput, latency, and cost trade-offs.
  • Experience navigating science-engineering tension — knowing when to hold the platform line and when to adapt for research velocity, and communicating that clearly to both sides.

Ideal candidate also has

  • Prior contributions to open source projects (Ray, Spark, RayDP, or Kubernetes).
  • Experience with enterprise security/compliance in ML environments.
  • Background working with edge/on-device ML and firmware/embedded teams.

Compensation and benefits

Annual base salary range: $196,000—$269,500 CAD. This role is also eligible for an initial RSU grant with no vesting cliff, and ongoing refresh opportunities tied to performance, subject to plan terms and conditions. Samsara offers a flexible, employee-led remote model, a professional development stipend, comprehensive health and parental leave plans, and more.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.