Skip to main content

Safeguards Enforcement Lead, User Well-Being

Anthropic
Hybrid - San Francisco, CA | New York City, NY | Washington, DC, at least 25% in officeUpdated 10d ago
Compensation
$285k–$330k
Published range · Top quartile for Operations (179 listings)
Location
Hybrid - San Francisco, CA | New York City, NY | Washington, DC, at least 25% in office
Remote eligibility
Employment
Full-time
Lead / Manager
Role family
Operations
AI / ML
Apply on job-boards.greenhouse.io
Job actionsApply now
Job actionsApply now

About the job

About the role

As a Safeguards Enforcement Lead on the User Well-Being team, you will manage child safety, mental health, abuse and exploitation, and age assurance enforcement workflows. This includes managing the team responsible for scaling, maintaining, and improving systems and processes to detect and respond to these harms. You will serve as the central point of contact for content review partners, managing day-to-day workflows, quality assurance, and escalation processes. You will also work closely with Engineering, Policy, and Legal teams to scale detection systems and close enforcement gaps.

Important context: In this role you will regularly be exposed to explicit content of a sexual nature involving minors, as well as content that may be violent or psychologically disturbing. Anthropic provides access to wellness resources and support.

Key responsibilities

  • Manage a team of individual contributors across multiple policy areas under the User Well-Being banner.
  • Serve as the primary point of contact for review partners, including onboarding, training, quality assurance, and relationship management.
  • Design and improve enforcement workflows to scale effectively while maintaining high accuracy and consistency.
  • Partner with Engineering and Data Science teams to optimize detection models and automated enforcement systems.
  • Develop and maintain internal documentation, decision trees, and review guidelines.
  • Keep up to date with emerging AI policy enforcement best practices and legal frameworks.
  • Identify and report trends in misuse patterns to internal stakeholders.
  • Coordinate reporting obligations to external bodies (e.g., NCMEC) in accordance with applicable law and Anthropic policy.

Minimum qualifications

  • Experience managing teams in the User Well-Being space.
  • Experience in trust & safety, content moderation operations, or policy enforcement with direct exposure to child safety, mental health, abuse and exploitation, and age assurance harm areas.
  • Experience managing or coordinating content review operations, including quality assurance and workflow management.
  • Experience standing up and scaling policy enforcement or content review workflows.
  • Proficiency in SQL and/or other data analysis tools to monitor workflow health and surface enforcement trends.
  • Experience identifying emerging risks and communicating findings to cross-functional stakeholders.
  • Understanding of challenges in implementing product policies at scale in content moderation.

Preferred qualifications

  • Deep subject matter expertise in child safety, CSEA, online child protection, mental wellness, and age assurance.
  • Experience working with or reporting to NCMEC, IWF, or equivalent.
  • Familiarity with legal frameworks including CSAM reporting obligations, KOSA, COPPA, or equivalent.
  • Experience with generative AI products and understanding of AI misuse.
  • Experience designing trauma-informed support structures and wellness protocols.
  • Proficiency in Python for workflow automation or data analysis.
  • Experience with hash-matching technologies (e.g., PhotoDNA, CSAI Match) or perceptual hashing tools.
  • Familiarity with age assurance technologies.

Compensation & benefits

Annual Salary: $285,000—$330,000 USD. Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space.

Application instructions

Apply via the Greenhouse job posting. Anthropic recruiters only contact from @anthropic.com addresses.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.