Automation QA w/ AI testing

Compunnel

$125K — $150K *
Finance & Insurance
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years in Quality Assurance, Quality Engineering, or Software Testing.
  • 2+ years evaluating AI/ML models or Generative AI testing.
  • Proficient in Selenium and Playwright for automation.
  • Solid understanding of AI/ML concepts and LLM testing methodologies.
  • Experience testing production-grade LLMs and AI applications.
  • Familiar with API testing and automation frameworks.
  • Strong analytical skills and problem-solving ability.

Responsibilities

  • Evaluate and validate AI/ML and Generative AI models for business and technical use cases.
  • Design robust evaluation strategies for model performance and accuracy.
  • Detect and address quality issues, including hallucinations, in AI outputs.
  • Develop automated and manual frameworks for testing AI models.
  • Collaborate with engineering, product, and data science teams on evaluation processes.
  • Utilize test automation tools for continuous integration and delivery.
  • Analyze the effectiveness of evaluation metrics such as precision and recall.

Benefits

  • Collaborative work environment with cross-functional teams.
  • Opportunities for using advanced AI technologies.
  • Involvement in innovative AI model evaluation projects.
Full Job Description
Job Summary

The Evaluation Engineer - AI Models will be responsible for evaluating and validating AI/ML and Generative AI models across business and technical use cases. This role combines extensive Quality Engineering and software testing experience with hands-on AI model evaluation, production-grade LLM testing, test automation, and data-driven analysis. The engineer will design evaluation strategies, validate model accuracy and performance, detect hallucinations and other quality issues, and develop automated and manual testing frameworks. Strong Selenium and Playwright expertise is required, along with excellent communication skills and the ability to collaborate with engineering, product, data science, and business stakeholders. Financial Services or Wealth Management domain experience is preferred.

Required Qualifications
• 10+ years of experience in Quality Assurance, Quality Engineering, Software Testing, or related disciplines.
• 2+ years of hands-on experience in AI model evaluation, Generative AI testing, or ML validation.
• Strong hands-on experience with Selenium and Playwright.
• Strong understanding of AI/ML concepts, LLM behavior, prompt evaluation, and model testing methodologies.
• Experience testing production-grade LLMs and AI-powered applications.
• Experience with API testing, test automation frameworks, and data validation techniques.
• Understanding of evaluation metrics such as precision, recall, accuracy, grounding, relevance, and hallucination detection.
• Experience creating automated and manual test strategies and frameworks.
• Strong analytical and problem-solving skills.
• Excellent verbal and written communication skills.
• Ability to work independently and collaborate effectively with cross-functional teams.

Similar Jobs

More Jobs at Compunnel

More Finance & Insurance Jobs

Find similar Automation QA w/ AI testing jobs: