The Effective Altruism
Opportunities Board
Work on the world's most pressing problems. Browse jobs, fellowships, internships, courses, and more at high-impact organisations.
Cyber Evaluations Engineer
AnthropicSan Francisco, CA / Washington, DC
San Francisco, CA / Washington, DC
Today
Routes to impact
Direct high impact on an important cause
Description
Build and run evaluations that measure cyber capabilities, misuse risks, and safeguard robustness in advanced AI models.
- Design cyber capability, uplift, and safety evaluations
- Test safeguards before major model releases
- Develop detection probes for cyber misuse
- Translate evaluation findings into safeguard improvements
This text was generated by AI. If you notice any inconsistencies, please let us know using this form.
Related opportunities
Full-Stack Software Engineer, Product
Apollo ResearchLondon, UK / San Francisco, CA
London, UK / San Francisco, CA
2 days ago
Software Engineer, Safeguards Evaluations
AnthropicSan Francisco, CA | New York City, NY
San Francisco, CA | New York City, NY
2 months ago
Machine Learning Engineer
10a LabsRemote (US)
Remote (US)
2 months ago
Security Engineer, Detection and Response, Japan
xAITokyo, Japan
Tokyo, Japan
3 months ago
Machine Learning Researcher
Gray SwanRemote
Remote
3 months ago
AI Security Research Engineer
0LabsRemote
Remote
4 months ago
Safeguards Enforcement Lead, Cyber Harms
AnthropicWashington, DC / San Francisco, CA / New York City, NY
Washington, DC / San Francisco, CA / New York City, NY
Yesterday
Software Engineer, Infrastructure, Interpretability
AnthropicSan Francisco, CA / New York City, NY
San Francisco, CA / New York City, NY
2 weeks ago
Join 60k subscribers and sign up for the EA Newsletter, a monthly email with the latest ideas and opportunities