Skip to main content

Researcher, Safety Training, National Security

OpenAI
San FranciscoUpdated 10d ago
Compensation
$380k–$500k
Published range · Top quartile for Data (97 listings)
Location
San Francisco
Remote eligibility
Employment
Full-time
Mid-level
Role family
Data
AI / ML
Apply on jobs.ashbyhq.com
Job actionsApply now
Job actionsApply now

About the job

About the Team

The Safety Training research team aims to fundamentally advance capabilities for precisely implementing safe behavior in AI models, and to leverage these advances to make OpenAI's deployed models safe and beneficial. Key focus areas include training nuanced safety behaviors, robustness to bad actors, addressing privacy and security risks, and trustworthiness in safety-critical situations.

About the Role

We're seeking a researcher to train and evaluate models for U.S. government use, with a focus on national security applications. You'll advance safety post-training and robustness, helping models follow nuanced policies while preserving usefulness and capabilities.

Responsibilities

  • Research and implement methods for safety training, reinforcement learning, and adversarial robustness.
  • Develop evaluations, identify model failure modes, and use findings to improve training.
  • Work with research, engineering, security, and policy partners to support safe, reliable deployment.

Qualifications

  • 4+ years of relevant AI safety research experience, including RLHF, adversarial training, or robustness.
  • Degree in computer science, machine learning, or related field, and strong deep learning research or engineering skills.
  • Experience improving model safety for deployment and enjoy collaborative research.
  • Motivated by OpenAI's mission and responsible use of AI in safety-critical settings.

Security Requirements

Active TS/SCI clearance or equivalent.

Compensation & Benefits

Salary range: $380K – $500K per year. Offers equity.

OpenAI is an equal opportunity employer and provides reasonable accommodations to applicants with disabilities.

Skills & tags

What you can verify before applying

Compare the essentials before you leave: pay, remote scope, employment type, source, and the employer apply destination.