Skip to main content

Software Engineer - Model Performance Systems

Baseten
Remote - San FranciscoUpdated 258d ago
Compensation
$165k–$330k
Published range
Location
Remote - San Francisco
Remote eligibility
Employment
Full-time
Mid-level
Role family
Engineering
AI / ML
Apply on jobs.ashbyhq.com
Job actionsApply now
Job actionsApply now

About the job

About Baseten

Baseten powers mission-critical inference for AI companies like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. They recently raised a $1.5B Series F led by Altimeter Capital, Conviction Partners, and Spark Capital.

The Role

This is a specialized, high-impact role at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will define the roadmap, drive key technical decisions, and take full ownership of the future of this work.

Responsibilities

  • Benchmarking: Evaluate, run, and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving).
  • DevEx Improvement: Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces).
  • Tool Development: Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking, and analysis.
  • System Profiling: Use profilers like PyTorch Profiler, NVIDIA Nsight Systems, and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack.
  • Monitoring & Observability: Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance.
  • Continuous Integration: Automate performance testing via CI/CD pipelines to catch regressions and build release workflow automation for the model runtimes stack.
  • Optimization Automation: Build tools to find the 'Pareto frontier'—identifying the best configuration (latency vs. cost vs. quality) for a given model and workload.

Requirements

  • This is a mid-senior, high leverage role. We care about technical depth, strong communication skills, ability to navigate vague requirements, and mentoring other engineers.
  • A love for systems and hardware: understanding GPU memory subsystems, InfiniBand, and how data moves across a cluster.
  • An automation mindset: if a task has to be done twice, it should be scripted; passion for stress-testing and fuzzy testing.
  • Mathematical curiosity: desire to understand the underlying math of Transformers and how it translates into FLOPs and memory requirements.
  • Technical toolkit: familiarity with Python, eagerness to master the NVIDIA software stack; C++ familiarity is good to have.

Benefits

  • Competitive compensation, including meaningful equity (U.S. only)
  • 100% coverage of medical, dental, and vision insurance for employee and dependents
  • Flexible PTO policy including company-wide Winter Break (offices closed from Christmas Eve to New Year's Day)
  • Paid parental leave
  • Fertility and family-building stipend through Carrot (U.S. only)
  • Company-facilitated 401(k)

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.