Skip to main content

Red Team Engineer, Safeguards

Anthropic
Hybrid - San Francisco, CA (at least 25% in office, travel required)Updated 30d ago
Compensation
$320k–$405k
Published range · Top quartile for Security (116 listings)
Location
Hybrid - San Francisco, CA (at least 25% in office, travel required)
Remote eligibility
Employment
Full-time
Mid-level
Role family
Security
AI / ML
Apply on job-boards.greenhouse.io
Job actionsApply now
Job actionsApply now

About the job

About the role

Anthropic's Safeguards team is seeking a Red Team Engineer to take an adversarial approach to uncover vulnerabilities across our product ecosystem before they can be exploited by malicious actors. The work spans technical infrastructure vulnerabilities to emergent risks from advanced AI capabilities, including coordinated account manipulation, payment fraud, and novel exploitation of product features.

Responsibilities

  • Conduct comprehensive adversarial testing across product surfaces, developing creative attack scenarios.
  • Research and implement novel testing approaches for emerging capabilities, including agent systems and tool use.
  • Design and execute 'full kill chain' attacks emulating real-world threat actors.
  • Build and maintain systematic testing methodologies and automated testing frameworks.
  • Collaborate with Product, Engineering, and Policy teams to translate findings into improvements.
  • Help establish metrics for measuring detection effectiveness.

Qualifications

Minimum: Experience in penetration testing, red teaming, or application security; model jailbreaking and testing large-scale agentic workflows for prompt injection vectors; strong web application security skills with tools like Burp Suite and Metasploit; building custom automation including LLM-specific testing frameworks; a track record of discovering novel attack vectors; a public body of work such as CVEs, blog posts, or bug bounty reports; strong communication skills.

Preferred: AI/ML security, understanding of AI safety, API security, business logic vulnerabilities, anti-fraud/trust & safety, distributed systems, and abuse detection.

Compensation

Annual salary: $320,000—$405,000 USD.

Benefits

Parental leave, flexible working hours, optional equity donation matching, generous vacation, and office space.

Application

Apply via the provided link.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.