About the Role
We are seeking a Lead AI/ML Engineer to architect, build, and operationalize a full-lifecycle, automated evaluation harness for non-deterministic Generative AI agents deployed across our Cloud GTM space (e.g., deep research, conversational analytics, pipeline forecasting).
In this role, you will transition our AI evaluations from manual human reviews to high-throughput automated pipelines backed by LLM-as-a-Judge and targeted human audits. You will work at the intersection of enterprise software engineering, distributed computing, and cutting-edge LLM evaluation frameworks to ensure data integrity, eliminate regressions, and scale our GTM agent capabilities.
Key Responsibilities
Foundations & Local Sandbox Development
Build robust golden datasets extracted from UAT logs and utilize frontier models to synthetically generate variations (typos, phrasing, syntax) for robust testing.
Production Pipeline Automation & System Architecture
Architect and deploy language-agnostic RPC endpoints to systematically invoke GTM agents within the ecosystem.
Build resilient production pipelines supporting parallel inference execution across 1,000+ trajectory datasets in under 15 minutes.
Stand up a centralized Model Context Protocol (MCP) logging server to capture raw prompts, tool trajectories, SQL queries, and token costs using strict JSON schemas.
Implement asynchronous message queues (Pub/Sub) for rate-limiting/backpressure, along with retry policies for network and generation failures.
Skill Benchmarking & Trajectory Validation
Author eval test suites to isolate specific agent competencies (e.g., CRM writes, SQL analytics).
Inject sandboxed mocks to validate tool-calling logic without producing live CRM side effects or executing heavy database reads.
Stream execution logs for baseline delta analysis and capability scoring.
Enterprise Analytics, Governance & Security (Phase 4)
Build low-latency Hydra ETL pipelines to stream structured evaluation JSON records into data warehouses and construct Plx analytics dashboards.