G2

Software Engineering Director, Agentic Evaluations

G2$150K — $180K *
US-AnywhereRemote in United States
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years in backend or full-stack programming
  • 2+ years of direct engineering management experience
  • Expert in backend programming languages such as Python, Java/Kotlin, or Go
  • Experience in creating evaluations for customer-facing agents
  • Familiarity with OpenAI, Anthropic, or Google LLMs
  • Daily use of coding agent tools like Codex or Pi
  • BS/BA degree in a relevant field

Responsibilities

  • Build and enhance evaluations for live agents
  • Scope feasibility for evaluating software vendor agents
  • Own the entire evaluation system's architecture
  • Design reusable evaluation primitives for multiple software categories
  • Establish repeatable processes for agent integration
  • Monitor new frameworks for evaluation improvements
  • Mentor and grow a team of engineers

Benefits

  • Flexible work environment
  • Opportunities for professional development
  • Collaborative team culture
  • Access to cutting-edge technology
  • Focus on work-life balance
Full Job Description
About The Role

Summary of Responsibilities:

As a Software Engineering Director, Agentic Evaluations, your responsibilities will center around delivery of useful and credible evaluations of agents offered by software vendors. You'll enhance the system that performs validation of agent performance, and lead the core team behind it. You will leverage "agent-first" ways of working, applying agentic engineering techniques to deliver your work to production.

Detailed Responsibilities:

Build and enhance agent evaluations that run against live agents [80%]:
  • Help scope out the feasibility and effort involved in evaluating an agent offered by a software vendor, identifying the most practical path to reaching a credible evaluation
  • Own the full stack of the core evaluation system, from admin & API surface to eval workflow dispatch
  • Design evaluation system primitives that generalize across software verticals, so onboarding a new category benefits from reuse
  • Distill the appropriate and repeatable processes in onboarding and maintaining integrations into repeatable AI skills or agents to increase evaluation velocity
  • Keep tabs on emerging frameworks and techniques for agent evaluation which offer improvements to evaluations, and provide industry-accepted means to distribute eval results
  • Collaborate with data science peers in their pursuit to develop rich and proprietary benchmarks

Lead a group of engineers and be a source of wider influence [20%]:
  • You will facilitate the growth and mentorship of a core team of engineers, and be a source of technical leadership and subject matter expertise in evals for them
  • You will also help socialize the use of evals in agent-oriented products built by the wider organization.
Minimum Qualifications:

We realize applying for jobs can feel daunting at times. Even if you don't check all the boxes in the job description, we encourage you to apply anyway.
  • 10+ years of professional programming experience in backend or full-stack environments
  • 2+ years of experience directly managing engineers
  • Expert-level proficiency in backend development using languages like Python, Java/Kotlin, Typescript/Javascript or Go; strong proficiency in associated backend frameworks like FastAPI or Node.js
  • Direct experience creating evaluations (evals) against a customer-facing agent, leveraging agent trajectory trace data and rubrics to measure an agent's task completion rate, accuracy, correctness, or policy adherence
  • Direct experience using frontier models from OpenAI, Anthropic, or Google in LLM-as-a-judge applications
  • Regular use of coding agent harnesses like Claude Code, Codex, Opencode, or Pi as part of the daily development workflow
  • BS/BA degree in related field
What Can Help Your Application Stand Out:
  • Experience with agent tool use via direct integration or as mediated by MCP servers
  • Knowledge of STATE-Bench, tau2-bench or similar benchmark frameworks
  • Experience using Playwright, browser-use, Chrome DevTools MCP or others to drive user flows in the browser
  • Experience with durable execution frameworks & workflow solutions, like Temporal, DBOS, Cloudflare / Vercel Workflows, or similar; or experience with agent sandboxing using AWS E2B, Daytona, Cloudflare / Vercel Containers or similar


About G2

G2 is a software company that provides a platform for business software reviews and ratings. The company was founded in 2012 and is headquartered in Chicago, Illinois. G2's platform allows users to read and write reviews of business software products, as well as compare products based on user ratings and other factors. The company has over 1 million reviews on its platform and serves over 5 million business professionals each month.
Learn more about G2
Industry
Founded
2012

Similar Jobs

More Jobs at G2

More Information Technology Jobs

Find similar Software Engineering Director, Agentic Evaluations jobs: