Skip to main content

Staff+ Software Engineer, Safeguards

Anthropic
Hybrid - San Francisco, CA or New York City, NY, at least 25% in officeUpdated 1d ago
Compensation
$320k–$485k
Published range · Top quartile for Engineering (554 listings)
Location
Hybrid - San Francisco, CA or New York City, NY, at least 25% in office
Remote eligibility
Employment
Full-time
Staff / Principal
Role family
Engineering
AI / ML
Apply on job-boards.greenhouse.io
Job actionsApply now
Job actionsApply now

About the job

Anthropic is a public benefit corporation building reliable, interpretable, and steerable AI systems. The Safeguards team builds safety and oversight mechanisms to monitor models, prevent misuse, and ensure user well-being, focusing on systems that detect unwanted model behaviors and prevent disallowed use of models while enforcing terms of service and acceptable use policies.

Multiple Safeguards teams are hiring; placement occurs after the interview process based on interests, experience, and organizational needs. Teams include Safeguards Acceleration (agentic systems extending Claude Tag for trust & safety work, sandboxed agent architectures, brokered data access, provenance guarantees), Safeguards Interventions (composable intervention options between the detection stack and users across 1P products, the API, and third-party clouds, including inline interventions for harmful bio, cyber, and acceptable usage), and Safeguards Data Intelligence (the Claude Investigation tool, an autonomous agent reasoning across billions of stored interactions to surface cyberattacks, weapons development, and state-sponsored influence operations, running an agent fleet across three clouds with batch inference, summarization, and search infrastructure).

Responsibilities

  • Develop monitoring systems to detect unwanted behaviors from API partners and take automated enforcement actions; surface in internal dashboards for manual review
  • Build abuse detection mechanisms and infrastructure
  • Surface abuse patterns to research teams to harden models at the training stage
  • Build robust multi-layered defenses for real-time improvement of safety mechanisms at scale

Qualifications

  • Bachelor's degree in Computer Science, Software Engineering, or comparable experience
  • Proficiency in Python and TypeScript; ability to work across the stack
  • Strong communication skills and ability to explain complex technical concepts to non-technical stakeholders
  • Strong candidates may also have 8+ years of software engineering experience; experience with integrity, spam, fraud, or abuse detection; trust and safety detection mechanisms for AI/ML systems; prompt engineering, jailbreak attacks, and adversarial inputs; or experience working with operational teams on custom internal tooling

Compensation

Annual salary: $320,000—$485,000 USD. Anthropic offers competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, and flexible working hours.

Application

No deadline; applications reviewed on a rolling basis. Visa sponsorship is available. Anthropic expects all staff to be in one of its offices at least 25% of the time.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.