BS/MS/PhD in Computer Science, Machine Learning, Statistics, or a related field (or equivalent experience).
At least 2 years of experience in ML evaluation, applied ML research, or a related engineering role.
Strong understanding of LLM fundamentals such as autoregressive generation and RLHF.
Proficiency in Python and ML frameworks like PyTorch.
Experience designing and implementing evaluation metrics for generative models.
Solid foundation in statistics and experimental design.
Experience with version control (Git) and containerization (Docker).
Responsibilities
Design and maintain evaluation frameworks for measuring LLM performance.
Define and implement quantitative metrics for model quality and reliability.
Build scalable, automated evaluation pipelines for model training and deployment.
Conduct statistical analysis to identify biases and performance gaps.
Translate real-world use cases into meaningful evaluation criteria.
Benefits
Work with world-class talent in AI research and development.
Opportunity to shape foundational technology.
Immediate impact on product direction and company trajectory.
Flexible vacation and paid time off (PTO).
Health, dental, and vision insurance.
Catered meals and commuter subsidies.
Collaborative and inclusive company culture.
Full Job Description
The Role
We seek experienced engineers and scientists to develop the evaluation metrics and systems that drive frontier LLM performance. You'll design the frameworks that tell us whether our models are improving and ensure they perform reliably at scale in production.
Key Responsibilities
Design, develop, and maintain robust evaluation frameworks and benchmarks for measuring LLM performance across diverse tasks and domains.
Define and implement quantitative metrics that capture model quality, safety, reliability, and regression detection.
Build scalable, automated evaluation pipelines that integrate into model training and deployment workflows.
Conduct rigorous statistical analysis of model outputs to identify failure modes, biases, and performance gaps.
Partner with product and customer-facing teams to translate real-world use cases into meaningful evaluation criteria.
Qualifications
BS/MS/PhD in Computer Science, Machine Learning, Statistics, or a related field (or equivalent experience).
At least 2 years of experience in ML evaluation, applied ML research, or a related engineering role.