Skip to main content

Safeguards Enforcement Analyst, Violence & Extremism

Anthropic
Hybrid - San Francisco, New York, or Washington, DC, at least 25% in officeUpdated 16d ago
Compensation
$285k–$330k
Published range · Top quartile for Operations (172 listings)
Location
Hybrid - San Francisco, New York, or Washington, DC, at least 25% in office
Remote eligibility
Employment
Full-time
Mid-level
Role family
Operations
AI / ML
Apply on job-boards.greenhouse.io
Job actionsApply now
Job actionsApply now

About the job

Anthropic is a public benefit corporation headquartered in San Francisco building reliable, interpretable, and steerable AI systems. As a Safeguards Enforcement Analyst focused on Violence & Extremism, you will build and execute operational workflows to assess model behavior, drive enforcement decisions, and develop evals across technically demanding policy areas, including weapons and dangerous technology, critical infrastructure attacks, violent extremism, and threats of violence. This role may involve exposure to explicit, violent, graphic, hateful, or psychologically disturbing content.

Key responsibilities

  • Design and architect automated enforcement systems and review workflows that scale while maintaining high accuracy
  • Develop and maintain evals that measure model performance on these policy areas, surface regressions, and inform policy and model improvements
  • Partner with Engineering and Data Science to optimize detection and automated enforcement systems for potential policy violations
  • Review flagged content to drive enforcement decisions and surface policy gaps, including novel or technically sophisticated misuse attempts and emerging extremist movements, ideologies, and mobilization tactics
  • Support the Safeguards policy design team with structured feedback on policy gaps and enforcement ambiguities
  • Develop and maintain enforcement guidelines and reviewer documentation
  • Keep up to date with emerging threats, terrorist and extremist movements, regulatory changes, and AI policy enforcement best practices
  • Identify and escalate emerging misuse patterns, novel attack vectors, and signs of coordinated violent extremist activity

Minimum qualifications

  • Experience in policy enforcement, threat intelligence, counterterrorism, government, or a closely related field, with direct exposure to harmful content, dangerous technology, violent extremism, or physical harm facilitation
  • Experience standing up and scaling policy enforcement or content review workflows
  • Proficiency in SQL and/or other data analysis tools to draw insights from large datasets and monitor enforcement workflow health
  • Experience identifying emerging risks and threat actors, and communicating findings to Product, Policy, Engineering, and Legal teams
  • Experience working with generative AI products, including writing effective prompts for content review and enforcement
  • Understanding of the challenges involved in implementing product policies at scale, including in the content moderation space

Preferred qualifications

  • Subject matter expertise in high-stakes harm areas such as weapons and dangerous technology, violent extremism, terrorism, autonomous systems, or critical infrastructure protection
  • Familiarity with relevant legal and regulatory frameworks governing dangerous technology, critical infrastructure, or domestic/international terrorism
  • Experience developing evals or red-teaming AI systems, particularly for harmful content or policy enforcement use cases
  • Experience with threat actor profiling and threat intelligence frameworks (e.g., MITRE ATT&CK)
  • Experience tracking threat actors, extremist networks, or misuse patterns across surface, deep, and dark web environments
  • Experience with large language models and an understanding of how AI technology could provide meaningful uplift toward serious harm
  • Proficiency in Python for data analysis and workflow automation
  • Background in law enforcement, national security, defense, counterterrorism, or a relevant regulatory environment
  • Familiarity with cross-platform threat analysis and OSINT techniques

Compensation

Annual salary: $285,000—$330,000 USD.

Logistics

Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience. Required field of study: a field relevant to the role. Location-based hybrid policy: all staff are expected to be in one of Anthropic's offices at least 25% of the time. Visa sponsorship is offered where possible.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.