AI JobFull-time
AI Researcher — Alignment & Safety
Anthropic
San Francisco, CA $220k–$350k Posted May 23, 2026
ai safetyinterpretabilityalignmentpytorchresearch
Job Description
Anthropic is seeking an AI Safety Researcher to advance our interpretability and alignment research programme. You will work on understanding how large language models represent information internally and designing training techniques that make models more reliably honest and safe.
Responsibilities
- Conduct original research on LLM interpretability and mechanistic analysis
- Design and run experiments to understand model behaviour at scale
- Develop alignment techniques that improve model honesty and corrigibility
- Write and publish research papers; present at top ML venues
Requirements
- PhD in ML, cognitive science, or related field preferred
- Track record of ML research publications
- Strong Python and PyTorch/JAX skills
- Deep curiosity about the mechanisms underlying LLM behaviour