Salary range: $250,000 - $500,000/year + benefits
About the role: We are looking for strong scientists and engineers to help advance our vision of oversight foundation models: AI systems that formalize, test, and answer questions about the behavior of other AI systems. As part of our highly collaborative team, you will learn and grow quickly, creating technology at the frontier of AI research and with high direct impact.
Core responsibility: Help us develop and train scalable oversight assistants that can predict and detect unexpected and subtle behaviors in AI systems. This includes:
- Creating diverse evaluations that range in difficulty. This involves finding naturally occurring interesting and undesirable behaviors exhibited by open-source models.
- Developing novel architectures and objectives for training oversight assistants.
- Scaling up the training and inference pipelines to support up to 1T-scale models.
Qualities of a strong candidate:- Experience with fine-tuning language models, designing new architectures, and creating evaluations.
- Reliable results: good experimental design, epistemic self-awareness and transparency
- Generativeness: coming up with original, productive ideas for unblocking progress
- Curiosity: a desire to understand ML systems and how they work
- Strong programming ability, including navigating trade-offs between prototyping speed and maintainability
- Strong communication skills, low ego, openness to giving and receiving feedback
We are located in San Francisco and enthusiastic to work together in-person. We are open to sponsoring international visas.