Research Engineer, Code RL (Reinforcement Learning)
Anthropic- Compensation
- $500k–$850k Published range · Top quartile for Engineering (520 listings)
- Location
- Hybrid - San Francisco, CA or New York City, NY, at least 25% in office Remote eligibility
- Employment
- Full-time Mid-level
About the job
About Anthropic
Anthropic is a public benefit corporation headquartered in San Francisco. Its mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial. The Reinforcement Learning teams have contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of the latest Claude models.
About the Role
As a Research Engineer on the Code RL team within the RL organization, you'll advance models' ability to write, edit, test, debug, and ship real software end to end on real codebases with real tools. The role blends research and engineering: design RL environments and coding tasks, build reward signals and verifiers that capture what 'good code' means, run training experiments on frontier models, diagnose why a model does or doesn't improve at a class of software-engineering work, and improve the speed and reliability of the pipelines. Focus areas span agentic coding behaviors, code correctness, long-horizon autonomous engineering, and high-performance code for accelerators.
Qualifications
Strong software-engineering skills and deep Python expertise, including async/concurrent programming; comfortable owning systems end to end and debugging across the stack; ability to balance research exploration with engineering implementation; care about code quality, testing, and performance; commitment to developing safe and beneficial AI systems. Strong candidates may also have experience with reinforcement learning, RLHF, post-training, or LLM finetuning; built coding agents, code-execution sandboxes, eval harnesses, verifiers, or developer tooling; background in program analysis, testing, verification, compilers, or formal methods; PyTorch and large-scale distributed training; CUDA/GPU/TPU kernel experience; virtualization and sandboxed code execution environments.
Compensation
Annual salary: $500,000—$850,000 USD.
Benefits
Competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space.
Application
Apply through the Greenhouse posting. Anthropic sponsors visas and will make every reasonable effort to obtain one for offered candidates. Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience.
Skills & tags
Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.