RoleYou'll define how frontier AI models are measured. You'll design new benchmarks, run experiments, analyze model behavior, and build evaluation methodologies that become trusted signals for the industry. Your work will shape our public leaderboards and the evaluation tools we share with frontier labs.
Here's an example of a piece of industry-leading work done in this field. This is a SOTA STS benchmark, advised by OpenAI: https://audioarena.ai/. Email us for the pre-print.
What You'll Own- Your work will be tracked and followed by the likes of Elon Musk, Mark Zuckerberg, Alexandr Wang, Demis Hassabis, Andrew Ng, Amjad Massad, and more
- Design genuinely hard and useful evaluations that measure frontier model performance on real-world tasks, and that become industry-leading gold-standards
- Investigate model failures and identify what they reveal about emerging capabilities
- Publish research, technical reports, and analyses that shape how frontier models are evaluated
What We're Looking For- Strong STEM background. You studied Computer Science, Data Science, Statistics, Math, Engineering, Physics, or a related field.
- Deep curiosity about frontier AI models. You're excited by understanding model behavior, discovering areas of failure, and building better ways to evaluate models.
- Genuine thirst and intellectual to be on the frontier of AI development.
- Fearlessness to roll up your sleeves, get your hands dirty, and do real work.
Details- Location: San Francisco, Levi's Plaza. We sponsor visas and handle relocation.
- Work Schedule: Sunday-Friday. Saturdays are yours!
- Compensation: Competitive salary + meaningful equity. You'd be joining at the stage when ownership matters most.