The Effective Altruism
Opportunities Board
Work on the world's most pressing problems. Browse jobs, fellowships, internships, courses, and more at high-impact organisations.
Researcher, Alignment Interpretability
OpenAISan Francisco, CA
San Francisco, CA
Today
Salary
$295K – $500K
Routes to impact
Direct high impact on an important cause
Description
Conduct mechanistic interpretability research to understand deep network representations and help ensure increasingly capable AI systems remain safe.
- Develop and publish techniques for understanding model representations
- Engineer infrastructure for studying model internals at scale
- Collaborate across teams on research projects
- Guide research toward usefulness and long-term scalability This text was generated by AI. If you notice any inconsistencies, please let us know using this form.
Related opportunities
Researcher, Recursive Self-Improvement Safety
OpenAISan Francisco, CA
San Francisco, CA
3 weeks ago
Researcher, Alignment Chain of Thought Monitorability
OpenAISan Francisco, CA
San Francisco, CA
4 weeks ago
AI Cyber Red Teamer
5 days ago
Senior Data Scientist, Safety
FacultyLondon, UK
London, UK
5 days ago
Research Scientist, Agent Robustness
ScaleSan Francisco, CA / New York, NY
San Francisco, CA / New York, NY
6 days ago
Member of Technical Staff, Embedded Assessments
Model Evaluation & Threat Research (METR)Berkeley, CA
Berkeley, CA
6 days ago
AI Security Researcher
MiniMaxBeijing, China / Shanghai, China
Beijing, China / Shanghai, China
1 week ago
Research Engineer - AI Verification
Singapore AI Safety Hub (SASH)Remote / London, UK / San Francisco, CA / Singapore
Remote / London, UK / San Francisco, CA / Singapore
2 weeks ago
Join 60k subscribers and sign up for the EA Newsletter, a monthly email with the latest ideas and opportunities