Software Engineer - Model Performance Systems
Baseten- Compensation
- $165k–$330k Published range
- Location
- Remote - San Francisco Remote eligibility
- Employment
- Full-time Mid-level
Role skills
Apply on jobs.ashbyhq.com
About the job
About Baseten
Baseten powers mission-critical inference for AI companies like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. They recently raised a $1.5B Series F led by Altimeter Capital, Conviction Partners, and Spark Capital.
The Role
This is a specialized, high-impact role at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will define the roadmap, drive key technical decisions, and take full ownership of the future of this work.
Responsibilities
- Benchmarking: Evaluate, run, and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving).
- DevEx Improvement: Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces).
- Tool Development: Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking, and analysis.
- System Profiling: Use profilers like PyTorch Profiler, NVIDIA Nsight Systems, and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack.
- Monitoring & Observability: Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance.
- Continuous Integration: Automate performance testing via CI/CD pipelines to catch regressions and build release workflow automation for the model runtimes stack.
- Optimization Automation: Build tools to find the 'Pareto frontier'—identifying the best configuration (latency vs. cost vs. quality) for a given model and workload.
Requirements
- This is a mid-senior, high leverage role. We care about technical depth, strong communication skills, ability to navigate vague requirements, and mentoring other engineers.
- A love for systems and hardware: understanding GPU memory subsystems, InfiniBand, and how data moves across a cluster.
- An automation mindset: if a task has to be done twice, it should be scripted; passion for stress-testing and fuzzy testing.
- Mathematical curiosity: desire to understand the underlying math of Transformers and how it translates into FLOPs and memory requirements.
- Technical toolkit: familiarity with Python, eagerness to master the NVIDIA software stack; C++ familiarity is good to have.
Benefits
- Competitive compensation, including meaningful equity (U.S. only)
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible PTO policy including company-wide Winter Break (offices closed from Christmas Eve to New Year's Day)
- Paid parental leave
- Fertility and family-building stipend through Carrot (U.S. only)
- Company-facilitated 401(k)
Skills & tags
What you can verify before applying
Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.