Staff+ Software Engineer, Safeguards Review Tooling
Anthropic- Compensation
- $320k–$485k Published range · Top quartile for Engineering (498 listings)
- Location
- Hybrid - San Francisco, CA, at least 25% in office Remote eligibility
- Employment
- Full-time Staff / Principal
About the job
Anthropic is a public benefit corporation headquartered in San Francisco, focused on creating reliable, interpretable, and steerable AI systems. The Safeguards team ensures models and products are developed and deployed safely.
About the role
The Review Tooling team builds the systems that humans — and increasingly Claude — use to investigate potential harms and take enforcement actions across Anthropic's first-party products and third-party cloud platforms. As one of the first engineers on this new team, you'll own the tools safety investigators rely on, as well as the platform underneath: analytics capabilities, privacy-preserving primitives, and a sandbox environment for building review interfaces. You'll also drive scaling review through automation, building systems where Claude extends what human reviewers can do while keeping people in the loop.
Key responsibilities
- Build investigation, review, and enforcement tooling for first-party and third-party platform surfaces, including case queues, investigation views, decision and audit logging, and account-actioning workflows.
- Develop the platform layer of reusable APIs, data storage, and backend services for quickly standing up new review workflows.
- Scale review through automation, including enabling reviewers to use Claude effectively and building toward Claude-assisted and Claude-driven review workflows.
- Partner with policy, operations, legal, privacy, and data science stakeholders to translate enforcement and investigation needs into reliable systems.
- Build guardrails: granular permissions, audit trails, data-access controls, and reviewer wellbeing features such as content obfuscation and exposure limits.
- Instrument shipped tools with metrics on queue health, reviewer throughput, and decision quality.
Minimum qualifications
- Technical background in full-stack or platform engineering.
- Experience shipping internal tools or platforms with demanding operational users.
- Experience working cross-functionally with non-engineering partners.
- Excellent communication skills.
- Care about the societal impacts of AI.
Preferred qualifications
- 8+ years of industry software engineering experience.
- Experience building trust and safety, integrity, fraud, or abuse-prevention tooling.
- Experience designing systems under strict privacy, compliance, or data governance constraints.
- Experience integrating LLMs or agentic systems into operational workflows, including agentic coding tools (e.g., Claude Code).
- Experience building developer platforms or extensible tooling frameworks.
- Experience supporting enforcement or moderation systems across multiple product surfaces.
Compensation and benefits
Annual salary: $320,000—$485,000 USD. Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a San Francisco office.
Logistics
Minimum education: Bachelor's degree or equivalent. Hybrid policy: all staff expected in one of Anthropic's offices at least 25% of the time. Visa sponsorship is offered.
Application
Apply through the Anthropic Greenhouse job posting.
Skills & tags
Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.