Key Responsibilities- The Benchmarking team defines how progress is measured. Researchers design evaluation frameworks that capture reasoning depth, interaction quality, reliability, and operational impact. They construct benchmarks that reflect real-world complexity. Their systems become the standard by which new architectures, techniques, and releases are judged.
- Researchers in Benchmarking explore new paradigms for evaluating intelligent systems: adversarial robustness testing, longitudinal performance tracking, and human-in-the-loop assessment. They investigate how metrics shape model behavior and establish rigorous methodologies for quantifying emergent capability. Their insights drive both Distyl's internal research priorities and industry-wide standards.
Who You Are- Experience Designing and Running Evaluations: You've built or maintained benchmarks, test suites, or experimental frameworks to measure model or system performance
- Statistical and Analytical Rigor: You design fair, reproducible experiments and can extract signal from noisy empirical results
- Experience Building with Models, Not Just Building Models: We develop intelligent systems using models rather than training or fine-tuning them. Ideal candidates have expertise in compound AI systems, agentic collaboration, and associated techniques (ensembling, ReAct, graph-of-thoughts, etc.)
- Proven Track Record of Research Results: Whether you've published in top journals, posted amazing work on twitter, or somewhere else we want to see what you've done
- Uses AI Every Day: Before you can revolutionize someone else's workflow, you need to revolutionize yours. You should be using tools like ChatGPT, Cursor, and Perplexity to accelerate your workflow
- Strong Programming and Data Analysis Skills: While you might not consider yourself a software engineer you need to be able to build prototypes of your ideas and then perform the experiments to prove the effectiveness to a F500 Head of AI
- Biases Towards Showing vs Telling: Our customers want to see the power of AI today vs discuss the most elegant idea that will take 5 years to realize
What We Offer- The base salary range for this role is $150K - $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package
- 100% coverage of medical, dental, and vision insurance for employee and dependents
- Flexible time off
- Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources
- Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot
- Complimentary in-office lunches and snacks provided
- Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems
- Ownership of high-impact projects across top enterprises
- A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence
Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday-Thursday) in-office.
#LI-Hybrid