Job Summary:
We are seeking a QA / Automation Engineer to build and maintain automated testing and evaluation capabilities for an agentic, LLM-driven application. The primary focus of this role is developing a new Python-based framework for evaluating agentic systems, including LLM-as-judge evaluators, conversational and agentic state validation, and other non-standard evaluation approaches that extend beyond conventional UI automation. The engineer will work collaboratively with developers in a pod-based environment where quality ownership is shared and the testing landscape is continuously evolving.
Key Responsibilities
• Build and maintain automated tests for agentic, LLM-driven applications beyond conventional UI test automation.
• Help develop a new Python-based framework for evaluating agentic systems.
• Design and implement LLM-as-judge style evaluators for assessing AI-driven behavior and responses.
• Develop state checks and behavior validation for conversational and agentic workflows.
• Explore and implement non-standard evaluation methods that do not map directly to traditional scripted automation.
• Use Playwright for front-end automation where applicable while maintaining a primary focus on Python-based AI evaluation.
• Collaborate closely with developers within a distributed QA and development pod structure.
• Contribute to quality ownership within a light, UAT-style QA process where testing responsibilities are shared across the development team.
• Operate effectively within an evolving testing landscape and contribute to building the framework as requirements and evaluation approaches develop.
Required Qualifications
• 7+ years of experience with Python, with Python serving as the primary language for the evaluation and automation framework.
• Experience testing AI/LLM-based systems, including evaluators, LLM-as-judge techniques, or state and behavior validation for agentic or conversational systems.
• Strong aptitude for developing testing approaches for emerging and evolving AI-driven systems.
• Working knowledge of Playwright for front-end automation.
• Ability to work effectively within a collaborative development and QA environment where quality ownership is shared with developers.
• Comfort working with a light, UAT-style formal QA process rather than a heavily siloed QA model.
• Ability to operate effectively in an undefined and continuously evolving testing landscape.
Preferred Qualifications
• HCM industry domain knowledge.
• Experience building or evaluating agentic AI and conversational applications.