Skip to main content

Research Engineer, Code RL (Reinforcement Learning)

Anthropic
Hybrid - San Francisco, CA or New York City, NY, at least 25% in officeUpdated 1d ago
Compensation
$500k–$850k
Published range · Top quartile for Engineering (520 listings)
Location
Hybrid - San Francisco, CA or New York City, NY, at least 25% in office
Remote eligibility
Employment
Full-time
Mid-level
Role family
Engineering
AI / ML
Apply on job-boards.greenhouse.io
Job actionsApply now
Job actionsApply now

About the job

About Anthropic

Anthropic is a public benefit corporation headquartered in San Francisco. Its mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial. The Reinforcement Learning teams have contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of the latest Claude models.

About the Role

As a Research Engineer on the Code RL team within the RL organization, you'll advance models' ability to write, edit, test, debug, and ship real software end to end on real codebases with real tools. The role blends research and engineering: design RL environments and coding tasks, build reward signals and verifiers that capture what 'good code' means, run training experiments on frontier models, diagnose why a model does or doesn't improve at a class of software-engineering work, and improve the speed and reliability of the pipelines. Focus areas span agentic coding behaviors, code correctness, long-horizon autonomous engineering, and high-performance code for accelerators.

Qualifications

Strong software-engineering skills and deep Python expertise, including async/concurrent programming; comfortable owning systems end to end and debugging across the stack; ability to balance research exploration with engineering implementation; care about code quality, testing, and performance; commitment to developing safe and beneficial AI systems. Strong candidates may also have experience with reinforcement learning, RLHF, post-training, or LLM finetuning; built coding agents, code-execution sandboxes, eval harnesses, verifiers, or developer tooling; background in program analysis, testing, verification, compilers, or formal methods; PyTorch and large-scale distributed training; CUDA/GPU/TPU kernel experience; virtualization and sandboxed code execution environments.

Compensation

Annual salary: $500,000—$850,000 USD.

Benefits

Competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and office space.

Application

Apply through the Greenhouse posting. Anthropic sponsors visas and will make every reasonable effort to obtain one for offered candidates. Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.