AI Evaluation Engineer (QA)

Appnovation Technologies

$100K — $120K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's Degree in a technical field or equivalent experience
  • 4+ years in QA/test engineering with a focus on data/ML systems
  • Proficient in Python and data-science techniques for quality measurement
  • Experience with LLM evaluation frameworks and statistical analysis
  • Skilled in creating automated, large-scale evaluation harnesses
  • Familiarity with output from various LLM providers
  • Strong background in test automation frameworks and scripting

Responsibilities

  • Run evaluations at scale across diverse question sets
  • Measure factual grounding and accuracy lift using statistical methods
  • Build metrics frameworks to demonstrate long-term quality improvement
  • Design tests to assess load and quality as data corpus expands
  • Define and maintain comprehensive test plans and quality gates
  • Automate regression and evaluation tests within CI/CD pipelines
  • Communicate quality metrics effectively to both technical and non-technical stakeholders
  • Collaborate with engineering teams to troubleshoot and verify fixes

Benefits

  • Opportunity to work on innovative QA strategies in AI
  • Collaborative and motivated team environment
  • Chance to influence quality standards and practices
  • Engagement with cutting-edge technology in ML and LLM
  • Development in a fast-paced, evolving field with growth opportunities
Full Job Description
As a QA / AI Evaluation Engineer, you will join a highly motivated and experienced team in a forward-leaning role that proves the platform actually improves answer quality. You will run evaluations at scale - from small human-UAT batches up to millions of automated evals - statistically measure factual grounding and accuracy lift, and build the metrics framework that shows how much better our answers get over time. We are looking for people who can bring a strong, solution-focused mindset and contribute to quality standards, best practices and get things done.
KEY RESPONSIBILITIES
  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvement (e.g., "answer is X% supported by source content / Y% better").
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.
QUALIFICATIONS
  • Bachelor's Degree in a technical field or equivalent experience.
  • 4+ years in QA / test engineering, with exposure to data/ML systems.
  • Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.
  • Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
  • Comfort working across multiple LLM providers' outputs.
  • Test automation frameworks and scripting.
  • Detail-oriented, with strong analytical and communication skills.
WHO YOU ARE
  • You think about how to scale, automate and operate, not just how to build a solution to an immediate problem
  • You understand lean thinking
  • You set high standards for code quality, performance/scalability and security and seek continuous improvement
  • You have solid analytical, problem solving and decision-making skills
  • You have customer first mindset and a devotion to customer service
  • You engage and build positive internal and external client relationships, while managing multiple initiatives, often with competing priorities
  • You have strong self-initiative, passion, interpersonal, oral and written communication and collaboration skills with the ability to work, influence and make an impact in a cross-functional environment with all levels of the organization
  • You are responsive and thrive in a fast-paced diverse high-performance environment with rapidly changing business needs
  • You actively seek out things outside your comfort zone with the ability to rapidly learn and take advantage of new concepts, business models, and technologies
  • You have prior experience in consulting
  • Prior experience and connections in the Life Sciences industry is preferred


Thank you for your interest in a career with Appnovation Technologies! Please note that only those selected for an interview will be contacted.

Similar Jobs

More Jobs at Appnovation Technologies

More Information Technology Jobs

Find similar AI Evaluation Engineer (QA) jobs: