[Expression of Interest] Research Engineer / Scientist, Alignment - London
Anthropic- Compensation
- £260k–£370k Published range
- Location
- Hybrid - London, UK, at least 25% in office Remote eligibility
- Employment
- Full-time Mid-level
About the job
Anthropic is a public benefit corporation focused on building reliable, interpretable, and steerable AI systems. The Alignment Science team conducts exploratory research on AI safety, with a focus on risks from powerful future systems (ASL-3/ASL-4).
As a Research Engineer/Scientist on the London team, you will design and run machine learning experiments to understand and steer AI behavior, collaborating with teams like Interpretability, Fine-Tuning, and the Frontier Red Team. Research areas include AI Control and Alignment Stress-testing.
Representative projects include testing safety techniques by training models to subvert them, running multi-agent reinforcement learning experiments (e.g., AI Debate), building tooling to evaluate LLM-generated jailbreaks, and contributing to research papers and talks.
You should have significant software, ML, or research engineering experience, some experience with empirical AI research, and familiarity with technical AI safety. Strong candidates may have authored research papers, experience with LLMs, reinforcement learning, or Kubernetes clusters.
Annual salary: £260,000—£370,000 GBP. Benefits include competitive compensation, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space.
This role is hybrid, requiring at least 25% time in the London office and occasional travel to San Francisco. Visa sponsorship is available.
Skills & tags
Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.