Full Job Description
QA Engineer, AI-Native Quality
Newton Research • Research & Development • Boston / Needham, MA
About the Role
Newton ships on a sprint cadence through a develop, stage and customer-environment pipeline, and the product surface is wide: conversations, blueprints, connectors, scheduled tasks, permissions and sharing, SSO, and AI agents whose behavior is not fully deterministic. A missed regression lands in front of a media planner or a customer's security review.
We run everything through an AI-first lens, because it is the only way quality scales. If a quality task is repeatable, an agent does it and you supervise; if it takes judgment, that is where you spend your time.
The gap this hire fills: evals and skill-change testing. Code changes already have CI and review, including PRs written by Claude. What has no safety net is behavior change: an edit to a skill, a prompt, a tool definition or a model version can silently change what our agents do, and nothing tests that today. You own that layer: curated eval sets, scoring and regression tracking, so any change to how an agent behaves is measured before it ships.
• Measure the AI: agent output varies run to run, so "correct" is a range you define with evals, rubrics and scoring, not an exact-match assertion.
• Let agents run the checks: agents run suites, triage failures and draft repro-ready defects; you design their roles and guardrails and review what they produce.
• Make Newton verifiable by agents: you keep the product and pipeline observable and fixture-friendly so agents can verify it without a human in the loop.
You are also a release-readiness partner (are we good?), alongside our existing QA lead, and that judgment stays human. But you are not hired to test Claude-driven PRs line by line.
What You Will Do
• Own the eval suite for Newton's agents: curated datasets of inputs and expected behaviors, rubric and LLM-as-judge scoring, regression tracking across model, prompt, skill and agent releases; validate judges against human-labeled examples
• Test skill and prompt changes before they ship: every edit to a skill, system prompt or tool definition runs against the relevant evals in CI, with a before/after comparison a reviewer can read in one glance; gate releases on the results
• Cover the agent flows end to end: tool-call correctness, task completion, multi-turn coherence, blueprint creation from conversations, code generation, scheduled tasks; flag output that is wrong, empty or silently degraded
• Handle non-determinism with rigor: repeated runs, pass-rate thresholds and simple statistics, so a flaky agent is a measured finding, not an anecdote
• Turn production and customer signal into evals: mine logs, error tracking and customer reports so every escaped bad-behavior case becomes a permanent eval, ideally drafted by an agent and reviewed by you
• Probe AI-specific risk: prompt injection, data leakage across users, projects and permissions, hallucinated or ungrounded numbers in analytics output, cost and latency regressions
• Automate before you repeat: any check you do by hand twice becomes a Playwright test, an eval or an agent workflow; drive manual regression time down every sprint
• Design agentic QA workflows in CI: agents that run suites, triage failures, draft defects and re-verify fixes, with guardrails, cost limits and escalation rules you define
• Keep human-judgment work sharp: exploratory testing across roles, feature flags and environments, SSO and connector authorization, and the cross-feature bugs only a person thinks to look for
• Write bug reports that are machine- and human-readable, and partner with engineering on testability (observability, seedable data, stable interfaces)
What Makes You a Great Fit
• 4+ years in QA or test engineering on a complex web product, ideally B2B SaaS shipping frequently
• Hands-on experience building evals for LLM or agent products (datasets, rubrics, LLM-as-judge, regression tracking), or a clear track record of getting there fast
• Working knowledge of how agents work: prompting, skills, tool calling, context, RAG and orchestration, well enough to tell where a failure originates
• An automation-first instinct: you can show what you removed from a manual process
• Python (or similar) for eval tooling and data checks; Playwright or similar for end-to-end tests
• Statistical literacy: pass rates, variance and sample size when outputs are not deterministic
• Daily user of AI coding and testing assistants, with the judgment to verify their output rather than trust it
• Strong exploratory instincts and excellent written communication
• Nice to have: adtech, martech or marketing analytics domain; SSO / SAML / OAuth flows; data-connector testing; eval or observability tooling
How Success Is Measured
• Eval suite covering Newton's core agent flows, run on every model, prompt, skill or agent change, with judge accuracy checked against human labels
• Share of skill and prompt changes that ship with a before/after eval result (target: all of them)
• Escaped bad-behavior cases converted into permanent evals
• Share of regression coverage run automatically or by agents, rising every sprint
• Release calls that hold up: few surprises after release
Salary range: $115,000-130,000 + Equity