In this role, you will:- Design, build, and scale evaluation frameworks for Zillow's agentic AI experiences.
- Build tracing, observability, and quality measurement systems for production AI agents.
- Create platform infrastructure that enables domain teams across Zillow to build on top of the Evals platform without bottlenecks.
- Partner with applied scientists and machine learning engineers to integrate new AI evaluation capabilities into production systems.
- Help evolve how Zillow measures AI quality, reliability, trustworthiness, and product impact.
- Stay current with emerging agentic AI paradigms, evaluation techniques, and LLM tooling, and translate them into practical platform innovation.
- Support scaling, reliability, performance optimization, incident response, and cost management for the evaluation layer.
- Apply first-principles thinking to ambiguous problems and iterate quickly on novel solutions.
This role has been categorized as a Remote position. "Remote" employees do not have a permanent corporate office workplace and, instead, work from a physical location of their choice, which must be identified to the Company. U.S. employees may live in any of the 50 United States, with limited exceptions.
In California, Connecticut, Maryland, Massachusetts, New Jersey, New York, Washington state, and Washington DC the standard base pay range for this role is $160,900.00 - $257,100.00 annually. This base pay range is specific to these locations and may not be applicable to other locations.In Colorado, Hawaii, Illinois, Maine, Minnesota, Nevada, Ohio, Rhode Island, Vermont, and Virginia the standard base pay range for this role is $152,900.00 - $244,300.00 annually. The base pay range is specific to these locations and may not be applicable to other locations.
In addition to a competitive base salary this position is also eligible for equity awards based on factors such as experience, performance and location. Actual amounts will vary depending on experience, performance and location. Employees in this role will not be paid below the salary threshold for exempt employees in the state where they reside.
You have:- 4+ years of backend engineering experience, with a track record of designing, shipping, and operating scalable production ML services
- Experience building platform infrastructure or developer-facing services consumed by multiple teams; you've thought carefully about APIs, user experience, reliability, and what it means to have internal customers
- Hands-on experience with ML Evals & Observability frameworks (Databricks MLflow, or evolving LLM frameworks preferred - LangSmith, Braintrust, Promptfoo, or similar)
- Experience evaluating third-party AI platforms and making principled build vs. buy decisions for platform infrastructure
- Fluency working across applied science, ML, product, design, and engineering; you translate between disciplines without losing precision
- Comfort with LLMs, agentic systems, and evaluation frameworks; familiarity with orchestration tools (LangChain, LangGraph, MCP, or similar)
- A habit of reaching for AI-assisted development tools in your own workflow; you don't just build for AI, you build with it
- Strong ownership and a growth mindset; you're energized by ambiguity, not slowed by it
Preferred qualifications:- Advanced degree in Computer Science, Machine Learning, or a related field, or equivalent practical experience building products using frontier LLMs, multimodal models, or agent-based systems
- Experience designing and operating evaluation pipelines: offline batch evals, online scoring, cost/latency/signal tradeoffs, and closing the loop from user signals back to training data.