AI Evaluation Engineer

Judi Health

$148K — $185K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4+ years of experience in data engineering, ML engineering, or software engineering
  • Bachelor's or Master's in Computer Science, Machine Learning, or a related field
  • Strong proficiency in Python for data manipulation and processing
  • Experience building and maintaining production data pipelines
  • Proficient in SQL and knowledgeable in cloud platforms (preferably AWS)

Responsibilities

  • Build and maintain ETL pipelines for various data sources
  • Develop dashboards and tools for monitoring AI quality metrics
  • Design metrics and evaluation methods for probabilistic systems
  • Partner with cross-functional teams to define evaluation success criteria
  • Implement automated evaluations and alerting on quality signals

Benefits

  • Hybrid work environment with flexibility to work remotely 3 days per week
  • Opportunity to work on cutting-edge AI evaluation and tooling
  • Engagement with cross-disciplinary teams to enhance AI model quality
  • Work in a dynamic, growth-focused company culture
  • Ability to influence the development of safety and quality benchmarks in AI systems
Full Job Description


Hybrid 3 days (offices in NYC, Denver, CO and Charlotte, NC area)

Position Summary

As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and autonomous agents in production. This role bridges the gap between model development and real-world usage by translating ambiguous product goals into measurable quality targets.

We're looking for someone to lead evaluation end-to-end - from unit and integration testing to offline, online, and statistical evaluations of probabilistic systems. What we need is someone who can design and operate robust evaluation frameworks, partner with scientists and engineers, and ensure we can confidently answer questions like: "Did this change improve or degrade quality, safety, or user outcomes?"

What You'll Build

Evaluation & Quality Pipelines
  • Build data evaluation pipelines that collect production conversations and agent interactions
  • Reconstruct full sessions from traces, logs, recordings, and transcripts
  • Apply labeling and scoring using human feedback signals (surveys, sentiment, outcomes) and automated evaluators (e.g., LLM-as-judge)

Continuous Quality & Safety Benchmarking
  • Own weekly and on-demand automated evaluation runs against staging and production
  • Define benchmarks that track accuracy, reliability, and safety-related signals
  • Produce trend dashboards that clearly answer: "Did this deploy change quality or risk?"

Unified Evaluation Framework
  • Design and extend a standardized evaluation framework that supports multiple agent types and workflows
  • Translate high-level product expectations into concrete success criteria and metrics
  • Ensure new agents and features can be evaluated consistently with minimal friction

Self Service Evaluation Tooling
  • Build APIs and internal tools so data scientists and engineers can go from "interesting scenario" to "included in the eval suite" quickly
  • Enable scenario curation, dataset management, and eval execution without deep infrastructure knowledge

Experiment Tracking & Visibility
  • Provide shared visibility into prompt, model, and agent experiments
  • Enable reproducibility and comparison across runs so teams can build on each other's work instead of operating in silos

Position Responsibilities:

Data Engineering
  • Build and maintain ETL pipelines for heterogeneous data sources (traces, logs, transcripts, user feedback)
  • Implement complex data stitching and session reconstruction logic
  • Manage dataset versioning, provenance, and lifecycle

Platform & Observability
  • Develop dashboards and monitoring tools for AI quality metrics
  • Integrate evaluations into CI/CD pipelines for scheduled and gated runs
  • Implement alerting on quality and safety signals, not just infrastructure health

AI / ML Evaluation Tooling
  • Apply and extend LLM-as-judge evaluation patterns
  • Design metrics and scoring approaches suitable for stochastic, non-deterministic systems
  • Use tools like LangSmith to track runs, traces, experiments, and evaluation results

Collaboration
  • Partner closely with data science, engineering, and product teams
  • Translate between research goals, product intent, and engineering constraints
  • Help define what "good" looks like for AI behavior in production
  • Advocate for strong developer experience and usability in the tools you build

Required Qualifications
  • 4+ years of experience in data engineering, ML engineering, or software engineering
  • Bachelor's or Master's degree in Computer Science, Machine Learning, or a related quantitative field
    • Strong proficiency in Python
    • Experience building and maintaining production data pipelines
    • Strong SQL skills
    • Experience working with at least one cloud platform (AWS preferred)

Nice-to-Haves
  • Prior work on LLM or agent evaluation infrastructure
  • Familiarity with designing metrics for safety, reliability, or quality in AI systems
  • Experience with voice or call-center data (audio, transcripts, sentiment)
  • Experience with browser automation tools (e.g., Playwright) for end-to-end evals
  • Deep SQL expertise


New York, NY Salary Range

$161,600-$200,000 USD

Denver, CO Salary Range

$148,400-$185,000 USD

Charlotte, NC Salary Range

$134,800-$168,500 USD

All employees are responsible for adherence to the Judi Health Code of Conduct including the reporting of non-compliance. This position description is designed to be flexible, allowing management the opportunity to assign or reassign duties and responsibilities as needed to best meet organizational goals.

Similar Jobs

More Jobs at Judi Health

More Information Technology Jobs

Find similar AI Evaluation Engineer jobs: