The Effective Altruism
Opportunities Board
Work on the world's most pressing problems. Browse jobs, fellowships, internships, courses, and more at high-impact organisations.
Researcher / PhD Student, Mechanistic Interpretability for Safe Agentic AI
Deutsches Forschungszentrum für Künstliche Intelligenz (DFKI) Saarbrücken, Germany
Saarbrücken, Germany
Today
Deadline
2026-09-15
Routes to impact
Direct high impact on an important cause
Skill-building & building career capital
Description
Advance mechanistic interpretability methods for controlling language models and improving the safety of agentic AI systems.
- Develop activation steering methods targeting specific computational circuits
- Study robustness and failure modes under distribution shifts
- Scale steering methods toward production-grade robustness
- Publish research at top-tier ML and NLP venues
This text was generated by AI. If you notice any inconsistencies, please let us know using this form.
Related opportunities
PhD Position, Responsible Machine Learning
ELLIS Institute TübingenVienna, Austria
Vienna, Austria
2 days ago
Research Engineer / Research Scientist, Misuse Red Team
AI Security Institute (AISI)London, UK
London, UK
1 week ago
Senior Researcher / Postdoc, AI Safety, Ethics, and Agentic Systems
Deutsches Forschungszentrum für Künstliche Intelligenz (DFKI) Saarbrücken, Germany / Darmstadt, Germany
Saarbrücken, Germany / Darmstadt, Germany
Today
Research Scientist, Control
Apollo ResearchSan Francisco, CA / London, UK
San Francisco, CA / London, UK
1 week ago
Research Scientist, Member of Technical Staff
AI DigestRemote
Remote
2 weeks ago
Member of Technical Staff, Engineering
AI DigestRemote
Remote
2 weeks ago
Research Assistant on AI Safety
University of OxfordOxford, UK
Oxford, UK
3 weeks ago
Researcher, Alignment Chain of Thought Monitorability
OpenAISan Francisco, CA
San Francisco, CA
1 month ago
Join 60k subscribers and sign up for the EA Newsletter, a monthly email with the latest ideas and opportunities