Responsibilities- Probe the model's internal representations for physical quantities, structure, and conservation laws
- Develop methods to explain individual predictions and the model's reasoning about interventions
- Investigate whether interventions in the model's internal state produce physically coherent responses
- Build tools and techniques for debugging model failures and understanding rollout behavior
- Partner with model, evaluation, and domain teams to turn interpretability findings into better models and greater trust
What we're looking forWe value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.
- Strong grasp of machine learning fundamentals and the internals of modern neural network architectures
- Experience or strong interest in interpretability, representation analysis, or related research
- Strong engineering skills for building interpretability tooling and running careful experiments
- A rigorous, hypothesis-driven approach to understanding model behavior
- A track record of turning open-ended research questions into concrete findings