Skip to main content

Software Engineer - GPU Networking & Distributed Systems

Baseten
Remote - San FranciscoUpdated 191d ago
Compensation
$165k–$330k
Published range
Location
Remote - San Francisco
Remote eligibility
Employment
Full-time
Mid-level
Role family
Engineering
AI / ML
Apply on jobs.ashbyhq.com
Job actionsApply now
Job actionsApply now

About the job

About Baseten

Baseten powers mission-critical inference for AI companies like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. The company recently raised a $1.5B Series F led by Altimeter Capital, Conviction Partners, and Spark Capital.

The Role

As a Software Engineer on the GPU Networking team, you will architect the software fabric that unifies thousands of GPUs into a cohesive operating system, making RDMA a first-class building block. You will go beyond network configuration to co-optimize communication and compute for Disaggregated Serving, Wide Expert Parallelism (WideEP), and fast cold starts.

Responsibilities

  • Integrate RDMA/RoCE/InfiniBand capabilities into the inference stack, moving beyond TCP/IP.
  • Implement and tune networking layers for Disaggregated KV Cache Offload and WideEP across NVLink and InfiniBand for MoE models.
  • Work with checkpointing and storage to enable sub-10-second startup for trillion-parameter models.
  • Characterize and validate networking performance on H100/H200, B200/B300, and GB200/300 NVL72 clusters.
  • Design tools to visualize packet flow, congestion, and effective bandwidth across GPU interconnects.
  • Work with communication libraries (NCCL, NVSHMEM) and potentially write custom communication kernels.

Requirements

  • Deep experience with high-performance networking protocols (InfiniBand, RoCE v2).
  • Fluent in C++ or Python.
  • Deep understanding of memory hierarchy in modern NVIDIA architectures (H100/Blackwell).
  • Comfort diving into TensorRT-LLM source code, writing custom C++/Python bindings, or debugging NVLink topology.

Nice to Have

  • Deep knowledge of NCCL, NVSHMEM, and UCX.
  • Experience with Rust for systems-level or performance-critical networking code.
  • Experience with GPUDirect Storage (GDS) or high-performance filesystems like Weka or 3FS.
  • Familiarity with TensorRT-LLM, vLLM, or SGLang.
  • Experience running low-level benchmarks to qualify new hardware clusters.

Compensation & Benefits

Salary range: $165K – $330K, plus equity. Benefits include 100% coverage of medical, dental, and vision insurance for employees and dependents, flexible PTO with a company-wide winter break, paid parental leave, fertility and family-building stipend through Carrot, and company-facilitated 401(k).

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.