QA Engineer-AI Native Quality

Newton Research

• $115K — $130K *
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4+ years in QA or test engineering, preferably in B2B SaaS
  • Experience building evals for LLM or agent products
  • Understanding of agent functionality and failure origins
  • Proficiency in Python or similar for eval tooling
  • Statistical literacy for analyzing non-deterministic outputs
  • Familiarity with AI coding and testing assistants
  • Strong exploratory instincts and excellent written communication

Responsibilities

  • Own the eval suite for agent behaviors and regression tracking
  • Test skill and prompt changes with before/after comparisons
  • Cover end-to-end agent flows and flag output issues
  • Handle non-determinism with rigorous statistical methods
  • Transform production signals into permanent evals
  • Probe AI-specific risks like prompt injection and data leakage
  • Automate manual checks to streamline regression testing
  • Design QA workflows that leverage agent capabilities
  • Conduct exploratory testing across various features
  • Write clear bug reports for both machines and humans

Benefits

  • Equity participation
  • Flexible work environment
  • Opportunities for professional growth
  • Collaborative team culture
  • Access to cutting-edge AI technologies
Full Job Description
QA Engineer, AI-Native Quality Newton Research • Research & Development • Boston / Needham, MA About the Role Newton ships on a sprint cadence through a develop, stage and customer-environment pipeline, and the product surface is wide: conversations, blueprints, connectors, scheduled tasks, permissions and sharing, SSO, and AI agents whose behavior is not fully deterministic. A missed regression lands in front of a media planner or a customer's security review. We run everything through an AI-first lens, because it is the only way quality scales. If a quality task is repeatable, an agent does it and you supervise; if it takes judgment, that is where you spend your time. The gap this hire fills: evals and skill-change testing. Code changes already have CI and review, including PRs written by Claude. What has no safety net is behavior change: an edit to a skill, a prompt, a tool definition or a model version can silently change what our agents do, and nothing tests that today. You own that layer: curated eval sets, scoring and regression tracking, so any change to how an agent behaves is measured before it ships. • Measure the AI: agent output varies run to run, so "correct" is a range you define with evals, rubrics and scoring, not an exact-match assertion. • Let agents run the checks: agents run suites, triage failures and draft repro-ready defects; you design their roles and guardrails and review what they produce. • Make Newton verifiable by agents: you keep the product and pipeline observable and fixture-friendly so agents can verify it without a human in the loop. You are also a release-readiness partner (are we good?), alongside our existing QA lead, and that judgment stays human. But you are not hired to test Claude-driven PRs line by line. What You Will Do • Own the eval suite for Newton's agents: curated datasets of inputs and expected behaviors, rubric and LLM-as-judge scoring, regression tracking across model, prompt, skill and agent releases; validate judges against human-labeled examples • Test skill and prompt changes before they ship: every edit to a skill, system prompt or tool definition runs against the relevant evals in CI, with a before/after comparison a reviewer can read in one glance; gate releases on the results • Cover the agent flows end to end: tool-call correctness, task completion, multi-turn coherence, blueprint creation from conversations, code generation, scheduled tasks; flag output that is wrong, empty or silently degraded • Handle non-determinism with rigor: repeated runs, pass-rate thresholds and simple statistics, so a flaky agent is a measured finding, not an anecdote • Turn production and customer signal into evals: mine logs, error tracking and customer reports so every escaped bad-behavior case becomes a permanent eval, ideally drafted by an agent and reviewed by you • Probe AI-specific risk: prompt injection, data leakage across users, projects and permissions, hallucinated or ungrounded numbers in analytics output, cost and latency regressions • Automate before you repeat: any check you do by hand twice becomes a Playwright test, an eval or an agent workflow; drive manual regression time down every sprint • Design agentic QA workflows in CI: agents that run suites, triage failures, draft defects and re-verify fixes, with guardrails, cost limits and escalation rules you define • Keep human-judgment work sharp: exploratory testing across roles, feature flags and environments, SSO and connector authorization, and the cross-feature bugs only a person thinks to look for • Write bug reports that are machine- and human-readable, and partner with engineering on testability (observability, seedable data, stable interfaces) What Makes You a Great Fit • 4+ years in QA or test engineering on a complex web product, ideally B2B SaaS shipping frequently • Hands-on experience building evals for LLM or agent products (datasets, rubrics, LLM-as-judge, regression tracking), or a clear track record of getting there fast • Working knowledge of how agents work: prompting, skills, tool calling, context, RAG and orchestration, well enough to tell where a failure originates • An automation-first instinct: you can show what you removed from a manual process • Python (or similar) for eval tooling and data checks; Playwright or similar for end-to-end tests • Statistical literacy: pass rates, variance and sample size when outputs are not deterministic • Daily user of AI coding and testing assistants, with the judgment to verify their output rather than trust it • Strong exploratory instincts and excellent written communication • Nice to have: adtech, martech or marketing analytics domain; SSO / SAML / OAuth flows; data-connector testing; eval or observability tooling How Success Is Measured • Eval suite covering Newton's core agent flows, run on every model, prompt, skill or agent change, with judge accuracy checked against human labels • Share of skill and prompt changes that ship with a before/after eval result (target: all of them) • Escaped bad-behavior cases converted into permanent evals • Share of regression coverage run automatically or by agents, rising every sprint • Release calls that hold up: few surprises after release Salary range: $115,000-130,000 + Equity

Similar Jobs

More Jobs at Newton Research

More Consumer Technology Jobs

Find similar QA Engineer-AI Native Quality jobs: