Safeguards Enforcement Analyst, Violence & Extremism
Anthropic- Compensation
- $285k–$330k Published range · Top quartile for Operations (172 listings)
- Location
- Hybrid - San Francisco, New York, or Washington, DC, at least 25% in office Remote eligibility
- Employment
- Full-time Mid-level
About the job
Anthropic is a public benefit corporation headquartered in San Francisco building reliable, interpretable, and steerable AI systems. As a Safeguards Enforcement Analyst focused on Violence & Extremism, you will build and execute operational workflows to assess model behavior, drive enforcement decisions, and develop evals across technically demanding policy areas, including weapons and dangerous technology, critical infrastructure attacks, violent extremism, and threats of violence. This role may involve exposure to explicit, violent, graphic, hateful, or psychologically disturbing content.
Key responsibilities
- Design and architect automated enforcement systems and review workflows that scale while maintaining high accuracy
- Develop and maintain evals that measure model performance on these policy areas, surface regressions, and inform policy and model improvements
- Partner with Engineering and Data Science to optimize detection and automated enforcement systems for potential policy violations
- Review flagged content to drive enforcement decisions and surface policy gaps, including novel or technically sophisticated misuse attempts and emerging extremist movements, ideologies, and mobilization tactics
- Support the Safeguards policy design team with structured feedback on policy gaps and enforcement ambiguities
- Develop and maintain enforcement guidelines and reviewer documentation
- Keep up to date with emerging threats, terrorist and extremist movements, regulatory changes, and AI policy enforcement best practices
- Identify and escalate emerging misuse patterns, novel attack vectors, and signs of coordinated violent extremist activity
Minimum qualifications
- Experience in policy enforcement, threat intelligence, counterterrorism, government, or a closely related field, with direct exposure to harmful content, dangerous technology, violent extremism, or physical harm facilitation
- Experience standing up and scaling policy enforcement or content review workflows
- Proficiency in SQL and/or other data analysis tools to draw insights from large datasets and monitor enforcement workflow health
- Experience identifying emerging risks and threat actors, and communicating findings to Product, Policy, Engineering, and Legal teams
- Experience working with generative AI products, including writing effective prompts for content review and enforcement
- Understanding of the challenges involved in implementing product policies at scale, including in the content moderation space
Preferred qualifications
- Subject matter expertise in high-stakes harm areas such as weapons and dangerous technology, violent extremism, terrorism, autonomous systems, or critical infrastructure protection
- Familiarity with relevant legal and regulatory frameworks governing dangerous technology, critical infrastructure, or domestic/international terrorism
- Experience developing evals or red-teaming AI systems, particularly for harmful content or policy enforcement use cases
- Experience with threat actor profiling and threat intelligence frameworks (e.g., MITRE ATT&CK)
- Experience tracking threat actors, extremist networks, or misuse patterns across surface, deep, and dark web environments
- Experience with large language models and an understanding of how AI technology could provide meaningful uplift toward serious harm
- Proficiency in Python for data analysis and workflow automation
- Background in law enforcement, national security, defense, counterterrorism, or a relevant regulatory environment
- Familiarity with cross-platform threat analysis and OSINT techniques
Compensation
Annual salary: $285,000—$330,000 USD.
Logistics
Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience. Required field of study: a field relevant to the role. Location-based hybrid policy: all staff are expected to be in one of Anthropic's offices at least 25% of the time. Visa sponsorship is offered where possible.
Skills & tags
Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.