Automation QA w/ AI testing

Compunnel

$110K — $130K *
Finance & Insurance
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years of experience in Quality Assurance or Quality Engineering.
  • 2+ years of hands-on experience in AI model evaluation and Generative AI testing.
  • Strong expertise in Selenium and Playwright for software testing.
  • Deep understanding of AI/ML concepts and model testing methodologies.
  • Experience with testing production-grade LLMs and AI applications.
  • Familiarity with API testing and data validation techniques.
  • Proficient in evaluation metrics such as precision and hallucination detection.

Responsibilities

  • Design evaluation strategies for AI/ML and Generative AI models.
  • Validate model accuracy and performance against business use cases.
  • Detect and report hallucinations and quality issues in AI outputs.
  • Develop automated and manual testing frameworks for model evaluation.
  • Collaborate with cross-functional teams including engineering and product stakeholders.
  • Perform data-driven analysis on testing results to inform decisions.
  • Evaluate and enhance existing testing methodologies and processes.

Benefits

  • Collaborative work environment with cross-functional teams.
  • Opportunity to work on cutting-edge technologies in AI/ML.
  • Exposure to projects in Financial Services and Wealth Management sectors.
  • Skill development through hands-on project experience.
  • Flexible work arrangements to promote work-life balance.
Full Job Description
Job Summary

The Evaluation Engineer - AI Models will be responsible for evaluating and validating AI/ML and Generative AI models across business and technical use cases. This role combines extensive Quality Engineering and software testing experience with hands-on AI model evaluation, production-grade LLM testing, test automation, and data-driven analysis. The engineer will design evaluation strategies, validate model accuracy and performance, detect hallucinations and other quality issues, and develop automated and manual testing frameworks. Strong Selenium and Playwright expertise is required, along with excellent communication skills and the ability to collaborate with engineering, product, data science, and business stakeholders. Financial Services or Wealth Management domain experience is preferred.

Required Qualifications
• 10+ years of experience in Quality Assurance, Quality Engineering, Software Testing, or related disciplines.
• 2+ years of hands-on experience in AI model evaluation, Generative AI testing, or ML validation.
• Strong hands-on experience with Selenium and Playwright.
• Strong understanding of AI/ML concepts, LLM behavior, prompt evaluation, and model testing methodologies.
• Experience testing production-grade LLMs and AI-powered applications.
• Experience with API testing, test automation frameworks, and data validation techniques.
• Understanding of evaluation metrics such as precision, recall, accuracy, grounding, relevance, and hallucination detection.
• Experience creating automated and manual test strategies and frameworks.
• Strong analytical and problem-solving skills.
• Excellent verbal and written communication skills.
• Ability to work independently and collaborate effectively with cross-functional teams.

Similar Jobs

More Jobs at Compunnel

More Finance & Insurance Jobs

Find similar Automation QA w/ AI testing jobs: