QA Engineer (AI Systems)

Nexxa.ai

$130K — $160K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in QA/SDET roles, owning test strategy for complex systems.
  • Hands-on experience testing LLM-based products and understanding generative system testing limitations.
  • Practical experience with evaluation frameworks for AI agents or building custom tools.
  • Strong scripting/programming skills, preferably in Python, for test automation and data pipeline development.
  • Familiarity with LLM-specific failure modes like hallucinations and prompt injections.
  • Ability to design and manage test data and labeled datasets over time.
  • Comfortable in ambiguous environments and defining quality metrics.
  • Strong communication skills for reporting quality findings to technical teams.

Responsibilities

  • Design evaluation harnesses and regression suites for LLM agents' components.
  • Develop datasets, including edge cases and adversarial prompts tailored for industrial contexts.
  • Track and define advanced quality metrics beyond basic accuracy.
  • Build automated testing pipelines integrated into CI/CD workflows.
  • Conduct adversarial testing with security teams to identify vulnerabilities.
  • Test agent behavior throughout the entire action loop, not just final outputs.
  • Investigate failures to identify the root cause across numerous system components.
  • Collaborate with engineers to convert evaluation failures into actionable bug reports.
  • Set quality standards for new agent capabilities before deployment.
  • Mentor engineers in testing strategies for probabilistic systems.

Benefits

  • Opportunity to work at the forefront of AI technology in industrial applications.
  • Collaborative work environment with cross-functional teams and experts.
  • Chance to mentor and shape the testing strategies of the team.
  • Involvement in building evaluation frameworks from the ground up.
  • Focus on innovative problem-solving in complex, real-world scenarios.
Full Job Description
Role Overview

We're looking for a Lead / Senior / Staff QA Engineer to own quality for Nexxa's AI agent systems - products that plan, call tools, and take multi-step actions autonomously in industrial environments. This isn't traditional UI testing: you'll be designing evaluation frameworks for non-deterministic, tool-using systems, building golden datasets, catching regressions in reasoning quality, and stress-testing agent behavior under adversarial and real-world edge-case conditions.

You'll work closely with ML engineers, backend engineers, and Forward Deployed Engineers to define what "good" looks like for an agent operating in high-stakes industrial settings, then build the infrastructure and processes to measure it continuously.

Key Responsibilities
  • Design and build evaluation harnesses and regression suites for LLM-based agents, covering reasoning quality, tool-call correctness, task completion, and multi-turn coherence.
  • Develop golden datasets and labeled test sets, including edge cases, ambiguous inputs, and adversarial prompts specific to industrial and operational contexts.
  • Define and track quality metrics beyond simple accuracy - groundedness, hallucination rate, task success rate, latency/cost tradeoffs, and safety violations.
  • Build automated pipelines that run evals on every model, prompt, or tool-integration change, and integrate them into CI/CD.
  • Conduct structured red-teaming and adversarial testing (prompt injection, jailbreaks, tool misuse, unsafe actions) in partnership with security teams.
  • Test agent behavior across the full action loop - planning, tool selection, tool execution, error recovery, and final output - not just the final response.
  • Investigate and triage failures where the root cause could be the model, the prompt, the tool/API, or the orchestration logic.
  • Partner with ML and backend engineers to translate eval failures into actionable, reproducible bug reports.
  • Establish quality bars and sign-off criteria for new agent capabilities before they reach customer environments.
  • Mentor other engineers on testing strategies specific to probabilistic, LLM-driven systems.
  • Advocate for testability and observability in agent architecture from day one.
Qualifications
  • 5+ years in QA/SDET roles, with demonstrated ownership of test strategy for complex systems.
  • Hands-on experience testing LLM-based products, chatbots, or AI agents - you understand why traditional deterministic test assertions break down for generative systems.
  • Practical experience with eval frameworks or tooling (e.g., promptfoo, DeepEval, RAGAS, LangSmith) or a track record of building your own.
  • Strong scripting/programming ability (Python preferred) to build test automation, data pipelines, and eval tooling.
  • Understanding of how LLM agents work: prompting, tool/function calling, context management, RAG, memory, and orchestration frameworks.
  • Experience designing test data and labeled datasets, including sourcing, sampling, and managing dataset drift over time.
  • Familiarity with LLM-specific failure modes: hallucination, prompt injection, context poisoning, tool misuse, goal drift, and non-determinism.
  • Comfortable operating in ambiguity - defining what "correct" means for a task when there's no single right answer.
  • Strong written communication skills for turning fuzzy quality signals into clear, actionable findings for engineering and product stakeholders.


Preferred
  • Experience with human-in-the-loop evaluation workflows (labeling pipelines, inter-rater reliability, rubric design).
  • Background in ML/data science sufficient to read model evals and statistical significance.
  • Experience red-teaming or doing adversarial/security testing on ML systems.
  • Familiarity with observability/tracing tools for LLM applications (e.g., LangSmith, Arize, Langfuse, Weights & Biases).
  • Experience testing AI systems in industrial, IoT, or operational technology (OT) environments.
  • Prior experience setting up eval infrastructure from scratch at a startup or fast-moving team.
  • What We're Looking For A QA engineer who wants to define what quality means for autonomous, real-world AI systems.
  • Someone who can build rigorous evaluation infrastructure for problems that don't have a single right answer.
  • A systems thinker who enjoys turning ambiguous agent behavior into measurable, trustworthy signals.
  • A strong collaborator who partners well with ML engineers, backend engineers, and Forward Deployed teams.

Similar Jobs

More Jobs at Nexxa.ai

  • AI Vision Engineer
    $130K — $160K *
    San Francisco, CA 94112 (San Francisco County)
    Technical Services
    In-Person
  • AI Vision Engineer
    $110K — $130K *
    Toronto, ON M3C 0E3
    Technical Services
    In-Person
  • QA Engineer (AI Systems)
    $130K — $160K *
    San Francisco, CA 94112 (San Francisco County)
    Enterprise Technology
    In-Person
  • Backend AI Engineer
    $150K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Information Technology
    In-Person
  • Staff DevOps Engineer
    $120K — $140K *
    Toronto, ON M3C 0E3
    Information Technology
    In-Person

More Enterprise Technology Jobs

Find similar QA Engineer (AI Systems) jobs: