Skip to main content

Research Engineer / Scientist, Alignment

Anthropic
Hybrid - San Francisco, CA, at least 25% in officeUpdated 1d ago
Compensation
$350k–$500k
Published range · Top quartile for Engineering (560 listings)
Location
Hybrid - San Francisco, CA, at least 25% in office
Remote eligibility
Employment
Full-time
Mid-level
Role family
Engineering
AI / ML
Apply on job-boards.greenhouse.io
Job actionsApply now
Job actionsApply now

About the job

About the role

Anthropic is a public benefit corporation headquartered in San Francisco, focused on creating reliable, interpretable, and steerable AI systems. As a Research Engineer on Alignment Science, you will contribute to exploratory experimental research on AI safety, focusing on risks from powerful future systems (ASL-3 or ASL-4 under the Responsible Scaling Policy), often collaborating with teams including Interpretability, Fine-Tuning, and the Frontier Red Team.

Responsibilities

  • Build and run elegant and thorough machine learning experiments to understand and steer the behavior of powerful AI systems.
  • Contribute to research areas including Scalable Oversight, AI Control, Alignment Stress-testing, Automated Alignment Research, Alignment Assessments, Safeguards Research, and Model Welfare.
  • Test robustness of safety techniques by training language models to subvert them.
  • Run multi-agent reinforcement learning experiments to test techniques like AI Debate.
  • Build tooling to evaluate novel LLM-generated jailbreaks.
  • Write scripts and prompts to produce evaluation questions for safety-relevant reasoning.
  • Contribute ideas, figures, and writing to research papers, blog posts, and talks.

Qualifications

  • Significant software, ML, or research engineering experience.
  • Some experience contributing to empirical AI research projects.
  • Familiarity with technical AI safety research.
  • Preference for fast-moving collaborative projects.
  • Strong candidates may have experience authoring research papers in ML, NLP, or AI safety; experience with LLMs, reinforcement learning, Kubernetes clusters, and complex shared codebases.

Compensation

Annual salary: $350,000—$500,000 USD.

Benefits

Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space.

Application instructions

Interviews are conducted in Python. Candidates should be based in the Bay Area. Visa sponsorship is available, though not for every role or candidate.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.