The Effective Altruism
Opportunities Board
Work on the world's most pressing problems. Browse jobs, fellowships, internships, courses, and more at high-impact organisations.
AI Security and Control Researcher
Apollo ResearchLondon, UK / San Francisco, CA
London, UK / San Francisco, CA
Today
Salary
San Francisco: $204,000 – $385,000; London: £136,000 – £258,000
Routes to impact
Direct high impact on an important cause
Description
Research coding agent security by developing threat models, failure modes, adversarial tests, and controls against potentially compromised or misaligned AI agents.
- Maintain security levels and a current library of coding agent failure modes
- Design realistic attacks for monitor development, evaluation, and backtesting
- Adjudicate flagged agent behavior and red-team monitoring systems
This text was generated by AI. If you notice any inconsistencies, please let us know using this form.
Related opportunities
AI Red Team Engineer
Apollo ResearchSan Francisco, CA / London, UK
San Francisco, CA / London, UK
Today
AI Security Researcher
Apollo ResearchLondon, UK / San Francisco, CA
London, UK / San Francisco, CA
2 days ago
Security Research Engineer
AI Verification and Evaluation Research Institute (AVERI)Remote / US / Mexico / Canada
Remote / US / Mexico / Canada
2 days ago
Member of Technical Staff, Embedded Assessments
Model Evaluation & Threat Research (METR)Berkeley, CA
Berkeley, CA
6 days ago
Member of Technical Staff, Cyberforensics
Model Evaluation & Threat Research (METR)Berkeley, CA
Berkeley, CA
6 days ago
Red Team Specialist, Cyber
OpenAISan Francisco, CA / Seattle, WA / Washington, DC
San Francisco, CA / Seattle, WA / Washington, DC
6 days ago
Research Scientist, Control
Apollo ResearchSan Francisco, CA / London, UK
San Francisco, CA / London, UK
Today
Product Security Engineer
Apollo ResearchSan Francisco, CA / London, UK
San Francisco, CA / London, UK
Today
Join 60k subscribers and sign up for the EA Newsletter, a monthly email with the latest ideas and opportunities