Minimum qualifications:- Bachelor's degree or equivalent practical experience.
- 8 years of experience in software development.
- 7 years of experience leading technical project strategy, ML design, and optimizing industry-scale ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
- 5 years of experience with design and architecture, and testing, launching software products.
- 2 years of experience with GenAI techniques (e.g., LLMs, Multi-Modal, Large Vision Models) or with GenAI-related concepts (language modeling, computer vision).
Preferred qualifications:- 10 years of experience leading the design and implementation of scalable distributed systems, ML platforms, or observability infrastructure, ideally in the context of large-scale agent platforms.
- Experience leading engineering teams and collaborating with product managers and researchers to translate scientific concepts into reliable developer tools.
- Strong understanding of LLM execution mechanics, tool-calling APIs, and model-level failure analysis.
- Fluency in distributed tracing, semantic monitoring, and telemetry collection at scale.
- Proven track record of shipping complex enterprise-grade developer platforms, virtualization/sandbox environments, or production-grade test harnesses.
About the jobWith your technical expertise you will manage project priorities, deadlines, and deliverables. You will design, develop, test, deploy, maintain, and enhance software solutions.
Google Cloud is building the most comprehensive AI platform in the industry. As platform adoption scales, enterprise customers require stable, highly scalable, and completely secure simulation environments and evaluation infrastructure to operationalize autonomous agents. This role focuses on the systems and platform engineering required to close the last mile of agentic governance for enterprise deployments.
You will lead the delivery of the engineering roadmap, integrating components from multiple product teams to build an intuitive governance user experience for CISOs, auditors, and enterprise IT teams.The Google Cloud AI Research team addresses AI challenges motivated by Google Cloud's mission of bringing AI to tech, healthcare, finance, retail and many other industries. We work on a range of unique problems focused on research topics that maximize scientific and real-world impact, aiming to push the state-of-the-art in AI and share findings with the broader research community. We also collaborate with product teams to bring innovations to real-world impact that benefits our customers. Individual pay is determined by factors including job-related skills, experience, and relevant education or training.
US: $262000 - $364000 (USD) 25% bonus target equity benefits
Learn more about benefits at Google .
Responsibilities- Architect, build, and operate large-scale environments that allow developers to test complex agent behaviors, tool calling, multi-turn workflows, and end-to-end workflows in isolated, reproducible sandboxes.
- Deliver reproducible evaluation infrastructure and automated test harnesses capable of validating against enteprise visibility, security, compliance requirements.
- Design and scale the platform infrastructure that enables agentic systems to validate online and audit offline and alert based on behavioral patterns.
- Build the metrics engines and dashboards to track evaluation coverage, semantic correctness, quality metrics, and regression signals over time. Standardize the use of metrics for production health.
- Integrate evaluation runs into deployment pipelines, creating release gates to prevent semantic and structural regressions before agent deployments reach production.
Information collected and processed as part of your Google Careers profile, and any job applications you choose to submit is subject to Google's Applicant and Candidate Privacy Policy .