About The RoleSummary of Responsibilities:As a Software Engineering Director, Agentic Evaluations, your responsibilities will center around delivery of useful and credible evaluations of agents offered by software vendors. You'll enhance the system that performs validation of agent performance, and lead the core team behind it. You will leverage "agent-first" ways of working, applying agentic engineering techniques to deliver your work to production.
Detailed Responsibilities:Build and enhance agent evaluations that run against live agents [80%]:- Help scope out the feasibility and effort involved in evaluating an agent offered by a software vendor, identifying the most practical path to reaching a credible evaluation
- Own the full stack of the core evaluation system, from admin & API surface to eval workflow dispatch
- Design evaluation system primitives that generalize across software verticals, so onboarding a new category benefits from reuse
- Distill the appropriate and repeatable processes in onboarding and maintaining integrations into repeatable AI skills or agents to increase evaluation velocity
- Keep tabs on emerging frameworks and techniques for agent evaluation which offer improvements to evaluations, and provide industry-accepted means to distribute eval results
- Collaborate with data science peers in their pursuit to develop rich and proprietary benchmarks
Lead a group of engineers and be a source of wider influence [20%]:- You will facilitate the growth and mentorship of a core team of engineers, and be a source of technical leadership and subject matter expertise in evals for them
- You will also help socialize the use of evals in agent-oriented products built by the wider organization.
Minimum Qualifications:We realize applying for jobs can feel daunting at times. Even if you don't check all the boxes in the job description, we encourage you to apply anyway.
- 10+ years of professional programming experience in backend or full-stack environments
- 2+ years of experience directly managing engineers
- Expert-level proficiency in backend development using languages like Python, Java/Kotlin, Typescript/Javascript or Go; strong proficiency in associated backend frameworks like FastAPI or Node.js
- Direct experience creating evaluations (evals) against a customer-facing agent, leveraging agent trajectory trace data and rubrics to measure an agent's task completion rate, accuracy, correctness, or policy adherence
- Direct experience using frontier models from OpenAI, Anthropic, or Google in LLM-as-a-judge applications
- Regular use of coding agent harnesses like Claude Code, Codex, Opencode, or Pi as part of the daily development workflow
- BS/BA degree in related field
What Can Help Your Application Stand Out:- Experience with agent tool use via direct integration or as mediated by MCP servers
- Knowledge of STATE-Bench, tau2-bench or similar benchmark frameworks
- Experience using Playwright, browser-use, Chrome DevTools MCP or others to drive user flows in the browser
- Experience with durable execution frameworks & workflow solutions, like Temporal, DBOS, Cloudflare / Vercel Workflows, or similar; or experience with agent sandboxing using AWS E2B, Daytona, Cloudflare / Vercel Containers or similar