AI Evaluation Engineer (QA)

Appnovation Technologies

$110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's Degree in a technical field or equivalent experience.
  • 4+ years in QA/test engineering, with experience in data/ML systems.
  • Strong Python skills for data science applications.
  • Experience with LLM evaluation frameworks and statistical analysis techniques.
  • Ability to automate large-scale evaluation harnesses and conduct human-in-the-loop tests.
  • Familiarity with multiple LLM providers.
  • Detail-oriented with strong analytical and communication skills.

Responsibilities

  • Run evaluations at scale across extensive question sets.
  • Statistically measure factual grounding and identify accuracy improvements.
  • Build a metrics framework to highlight quality enhancements over time.
  • Design load and quality tests as the corpus expands.
  • Define and manage test plans, test cases, and quality gates.
  • Automate regression and evaluation suites and integrate them into CI/CD processes.
  • Report quality metrics clearly to both technical and non-technical stakeholders.
  • Collaborate with engineering teams on reproducing, triaging, and verifying fixes.

Benefits

  • Engaging work environment with a focus on continuous improvement.
  • Opportunity to work with advanced AI technologies.
  • Collaborative culture with cross-functional initiatives.
  • Support from a motivated and experienced team.
Full Job Description
As a QA / AI Evaluation Engineer, you will join a highly motivated and experienced team in a forward-leaning role that proves the platform actually improves answer quality. You will run evaluations at scale - from small human-UAT batches up to millions of automated evals - statistically measure factual grounding and accuracy lift, and build the metrics framework that shows how much better our answers get over time. We are looking for people who can bring a strong, solution-focused mindset and contribute to quality standards, best practices and get things done.
KEY RESPONSIBILITIES
  • Run evals at scale across large question sets, from small human-UAT batches up to hundreds of thousands or millions of automated evaluations.
  • Statistically measure factual grounding and accuracy lift (before/after), not just human A/B testing.
  • Build a metrics framework showing quality improvement (e.g., "answer is X% supported by source content / Y% better").
  • Design load and quality tests as the corpus scales.
  • Define and maintain test plans, test cases, and quality gates.
  • Automate regression and evaluation suites; integrate them into CI/CD.
  • Report quality metrics clearly to technical and non-technical stakeholders.
  • Collaborate with engineering to reproduce, triage, and verify fixes.
  • Continuously improve QA processes and coverage.
QUALIFICATIONS
  • Bachelor's Degree in a technical field or equivalent experience.
  • 4+ years in QA / test engineering, with exposure to data/ML systems.
  • Strong Python and data-science techniques for measuring factual grounding and answer quality.
  • Experience with LLM evaluation frameworks and statistical analysis.
  • Ability to build automated, large-scale eval harnesses plus lighter human-in-the-loop A/B tests.
  • Comfort working across multiple LLM providers' outputs.
  • Test automation frameworks and scripting.
  • Detail-oriented, with strong analytical and communication skills.
WHO YOU ARE
  • You think about how to scale, automate and operate, not just how to build a solution to an immediate problem
  • You understand lean thinking
  • You set high standards for code quality, performance/scalability and security and seek continuous improvement
  • You have solid analytical, problem solving and decision-making skills
  • You have customer first mindset and a devotion to customer service
  • You engage and build positive internal and external client relationships, while managing multiple initiatives, often with competing priorities
  • You have strong self-initiative, passion, interpersonal, oral and written communication and collaboration skills with the ability to work, influence and make an impact in a cross-functional environment with all levels of the organization
  • You are responsive and thrive in a fast-paced diverse high-performance environment with rapidly changing business needs
  • You actively seek out things outside your comfort zone with the ability to rapidly learn and take advantage of new concepts, business models, and technologies
  • You have prior experience in consulting
  • Prior experience and connections in the Life Sciences industry is preferred


Thank you for your interest in a career with Appnovation Technologies! Please note that only those selected for an interview will be contacted.

Similar Jobs

More Jobs at Appnovation Technologies

More Information Technology Jobs

Find similar AI Evaluation Engineer (QA) jobs: