Skip to main content

Research Engineer, RL Scaling Science

Anthropic
Hybrid - London, UK, at least 25% in officeUpdated 1d ago
Compensation
£375k–£640k
Published range
Location
Hybrid - London, UK, at least 25% in office
Remote eligibility
Employment
Full-time
Mid-level
Role family
Engineering
AI / ML
Apply on job-boards.greenhouse.io
Job actionsApply now
Job actionsApply now

About the job

About the role

Anthropic's RL Scaling Science team studies how reinforcement learning behaves as it scales (across model size, compute, and task horizon) and turns that understanding into the training recipes behind frontier models. As a Research Engineer, you'll design and run large-scale experiments to understand and resolve bottlenecks, build benchmarks that make long-horizon progress measurable, and ship validated findings directly into production training.

Key responsibilities

  • Design, run, and interpret large-scale RL experiments, reasoning rigorously about what the data does and doesn't show
  • Investigate how RL improves as horizon, compute, and model size grow
  • Build and maintain benchmarks for long-horizon RL so progress is measurable and reproducible
  • Translate validated findings into production training recipes, exercising judgment about when a result is robust enough to ship
  • Debug complex issues at the seam where research meets infrastructure - failures that only appear at scale
  • Partner closely with adjacent RL teams across research and engineering and advance the overall RL stack

Minimum qualifications

  • Strong empirical research skills in Reinforcement Learning, large-scale ML training, or a closely adjacent area
  • Demonstrated ability to own large experiments end-to-end, from design through interpretation
  • Proficiency in Python and experience working with large-scale or distributed ML systems
  • Comfort operating at the research/systems boundary, including debugging where the two meet
  • Care about the societal impacts of AI and responsible scaling

Preferred qualifications

  • Published or shipped work in long-horizon RL or RL fundamentals
  • Experience translating research findings into production training recipes
  • Demonstrated large scale industry impact via RL interventions
  • Experience working on frontier-scale training runs with long trajectories

Compensation

Annual Salary: £375,000—£640,000 GBP

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.