Responsibilities- Own the quality strategy for teams shipping largely agent-generated code, ensuring software can be released safely and at speed through strong verification practices.
- Build and maintain automated test infrastructure across unit, integration, end-to-end, contract, API, performance, and regression testing.
- Define and operate evaluation frameworks for AI-powered features, including datasets, scoring systems, model-as-judge approaches, human calibration, regression tracking, and release gates.
- Design testing strategies for non-deterministic and agentic systems where outputs may vary while still meeting expected quality standards.
- Use AI coding agents to generate and maintain test suites, while ensuring coverage remains meaningful, risk-based, and aligned to business requirements.
- Test agent behavior, including multi-step task execution, tool usage, failure recovery, safety enforcement, and human handoff workflows.
- Instrument production quality signals and use real-world feedback to continuously improve evaluation and testing frameworks.
- Partner closely with engineering teams throughout the development lifecycle, embedding quality from refinement through release.
Requirements- 8+ years of industry experience in test automation or quality engineering, with strong programming skills in Python, TypeScript, Java, or C# and a track record of supporting complex software products.
- Hands-on experience testing AI and model-based systems, including evaluation frameworks, model quality assessment, and the challenges of non-deterministic outputs.
- Proven ability to build and scale automated testing platforms, including CI/CD integration, test data strategies, framework development, and flaky test management.
- Strong understanding of how to evaluate agentic systems, including tool calling, multi-step workflows, safety mechanisms, and failure handling.
- Regular use of AI coding agents, combined with a critical review mindset and accountability for the correctness and quality of generated software and test assets.
- Experience with API testing, performance testing, and at least one of desktop, web, or distributed backend platforms.
- Deep understanding of common AI failure modes, including hallucinations, prompt injection, context degradation, non-determinism, and silent regressions, along with effective mitigation strategies.
- Practical experience with agentic development environments, MCP or comparable tooling, and repository-based management of prompts, instructions, and evaluation assets.
- Strong communication skills, the ability to collaborate effectively across distributed teams, and a passion for helping shape a modern AI-native.
What we offer- Competitive compensation and bonuses
- Flexible PTO and paid holidays
- 401(k) with employer matching
- Comprehensive Health insurance package including 100% employer-paid medical coverage
- Up to 12 weeks of Parental Leave
- Basic Life Insurance, Short-Term & Long-Term Disability, 100% employer-paid
- Quarterly teambuilding events, leadership luncheons, and companywide "All Hands" meetings
- Open door policy and business casual dress code
- We celebrate diversity as one of our core values. Join c-a-r-e and lead change initiatives together with us!
Work location for this position is Austin, TX.
Department Research & Development Locations Austin Remote status Hybrid Employment type Full-time Type of Job Non Student