The Effective Altruism
Opportunities Board
Work on the world's most pressing problems. Browse jobs, fellowships, internships, courses, and more at high-impact organisations.
Model Policy Manager, Agentic Safety
OpenAISan Francisco, CA
San Francisco, CA
Today
Salary
$207K – $335K
Routes to impact
Direct high impact on an important cause
Skill-building & building career capital
Description
Shape behavioral safety policy for frontier AI systems by translating alignment risks into safeguards, evaluations, and monitoring.
- Investigate model misalignment across long-horizon agent behavior.
- Develop policies, threat models, evaluations, and safety safeguards.
- Partner with research, engineering, security, and product teams.
- Inform deployment decisions and post-deployment monitoring.
This text was generated by AI. If you notice any inconsistencies, please let us know using this form.
Related opportunities
Research Engineer, Human Influence
AI Security Institute (AISI)London, UK
London, UK
1 day ago
Research Resident, Emerging Technology and Security, Open Source
RANDRemote (U.S.)
Remote (U.S.)
1 month ago
Senior Research Scientist, AI Safety Evaluations
FacultyLondon, UK
London, UK
1 month ago
Safeguards Enforcement Analyst, Violence and Extremism
AnthropicSan Francisco, CA / New York City, NY / Washington, DC
San Francisco, CA / New York City, NY / Washington, DC
1 month ago
Risk Modeling Lead
SaferAIParis, France / London, UK / San Francisco, CA / Remote
Paris, France / London, UK / San Francisco, CA / Remote
2 months ago
Research Scientist/Engineer (Science of Scheming)
Apollo ResearchLondon, United Kingdom
London, United Kingdom
6 months ago
Research Scientist/Engineer (Evaluations)
Apollo ResearchLondon, United Kingdom
London, United Kingdom
6 months ago
Data Scientist, Cybersecurity
OpenAIRemote (New York City, NY / San Francisco, CA)
Remote (New York City, NY / San Francisco, CA)
4 weeks ago
Join 60k subscribers and sign up for the EA Newsletter, a monthly email with the latest ideas and opportunities