Senior Software Engineer, Machine Learning Platform
Chime- Base salary
- $187k–$259k Published base salary range
- Location
- Hybrid - San Francisco, CA, USA, 4 days/week in office Remote eligibility
- Employment
- Full-time Senior
About the job
About the role
Chime’s Machine Learning Platform (MLP) team builds and operates the infrastructure, tooling, and developer experience that powers machine learning across the company. As a Senior Software Engineer, you will design and build scalable systems spanning traditional ML and emerging AI workloads, including model training, feature computation, real-time inference, foundation-model access, evaluation, and agentic orchestration.
What you can expect
- Design, build, and operate scalable ML and AI infrastructure on AWS.
- Design and operate shared platform capabilities for LLM and agentic workloads, including model access, prompt and configuration lifecycle, retrieval, tool integration, state management, and workflow orchestration.
- Build evaluation frameworks for non-deterministic AI systems, including offline benchmarks, regression testing, online quality signals, human feedback, and failure analysis.
- Establish observability, reliability, and governance for models and agents, covering traces, model and prompt versions, tool calls, latency, token usage, quality, safety, privacy, and cost.
- Build distributed training, batch inference, and large-scale processing systems using frameworks such as Ray or Spark.
- Build and maintain infrastructure as code using Terraform.
- Develop data ingestion and streaming systems using technologies such as Kinesis, Kafka, Flink, or Spark.
- Improve CI/CD workflows for ML models, AI applications, and platform components.
- Participate in on-call rotations to support production systems.
To thrive in this role, you have
- Knowledge of the machine learning development lifecycle, including data preprocessing, model training, evaluation, deployment, and monitoring.
- Experience designing distributed systems and large-scale data or compute platforms on AWS using frameworks such as Spark or Ray.
- 5+ years of experience in ML or AI infrastructure, platform engineering, distributed systems, or production ML systems.
- Working knowledge of LLM application patterns such as retrieval-augmented generation, structured outputs, tool calling, agent orchestration, and evaluation of non-deterministic systems.
- Hands-on experience with CI/CD pipelines, DevOps practices, and infrastructure as code.
- Experience with containerization and orchestration technologies such as Docker and Kubernetes.
- Strong programming skills in Python, Go, Scala, Java, or similar languages.
Nice-to-have
- Experience shipping LLM-powered or agentic systems to production.
- Experience with model gateways, prompt lifecycle management, retrieval or vector search, tool execution, and agent orchestration frameworks.
- Familiarity with managed or self-hosted foundation model infrastructure, such as Amazon Bedrock, SageMaker, or equivalent platforms.
- Experience operating GPU-based workloads and optimizing training or inference performance and cost; CUDA experience is a plus.
Compensation & benefits
Base salary offered for this role and level of experience: $187,000—$259,000 USD. Full-time employees may also be eligible for bonus(es), competitive equity, and benefits. Chime offers comprehensive health, financial, and wellbeing benefits, generous vacation policy and company-wide paid days off, annual wellness stipend, up to 22 weeks of paid parental leave for birthing parents and 12 weeks for non-birthing parents, family planning reimbursement, and more.
Application instructions
Apply via the Greenhouse job posting. For accommodation during the application process, contact accommodations@chime.com.
Skills & tags
Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.