Research Engineer, Performance RL (Reinforcement Learning)
Anthropic- Compensation
- $350k–$850k Published range · Top quartile for Engineering (597 listings)
- Location
- Hybrid - San Francisco, CA, at least 25% in office Remote eligibility
- Employment
- Full-time Mid-level
About the job
Anthropic is a public benefit corporation headquartered in San Francisco, focused on creating reliable, interpretable, and steerable AI systems. The Reinforcement Learning (RL) teams lead Anthropic's RL research and development, contributing to all Claude models, including significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.6 and Opus 4.6. The Code RL team within the RL organization is hiring a Research Engineer to advance models' ability to safely write correct, fast code for accelerators.
Responsibilities
- Invent, design, and implement RL environments and evaluations.
- Conduct experiments and shape the research roadmap.
- Deliver work into training runs.
- Collaborate with researchers, engineers, and performance engineering specialists.
Qualifications
- Expertise with accelerators (CUDA, ROCm, Triton, Pallas) and ML framework programming (JAX or PyTorch).
- Experience across the stack – kernels, model code, distributed systems.
- Ability to balance research exploration with engineering implementation.
- Passion for AI's potential and commitment to safe and beneficial systems.
- Strong candidates may have experience with reinforcement learning, porting ML workloads between accelerators, and familiarity with LLM training methodologies.
Compensation: Annual salary range $350,000—$850,000 USD.
Logistics: Minimum education: Bachelor's degree or equivalent. Location-based hybrid policy: all staff expected in one of our offices at least 25% of the time. Visa sponsorship is available.
Skills & tags
Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.