OverviewHow do you know an AI product is actually getting better-and how do you prove it at scale, before millions of people experience the difference?
We build and operate the offline evaluation platform that gates Microsoft Copilot's quality. Teams across Copilot depend on us to run their scenarios against the product, score the responses, and produce the scorecards that inform what ships. We do this at the scale of one of the world's largest AI products, within a strict enterprise compliance boundary. As Copilot evolves into a fleet of autonomous agents, the way we measure quality must evolve with it-and you will lead that transformation.
As a
Principal Software Engineer on the Copilot evaluation team, you will own the technical vision for a platform built AI-first and agent-first: one where autonomous agents are not only measured, but also help operate, extend, and heal the system itself.
You will take on our hardest problem: delivering trustworthy results across a deep chain of dependencies-including Copilot's model, retrieval, and scoring services-that we push far beyond their normal operating envelope, with hundreds of calls per job under relentless and growing load. You will architect how we evaluate agentic Copilot experiences, from multi-turn trajectories and tool selection to end-to-end task outcomes, and set the standard for how the broader product defines and measures quality.
This opportunity will allow you to shape the evaluation strategy for one of the world's most visible AI products, work at the frontier of large-scale agentic systems and reliability engineering, and grow your influence as a technical leader across a large engineering organization.
Responsibilities- Sets the technical direction for the Copilot offline evaluation platform and partners with teams across the Copilot organization to turn ambiguous quality questions into rigorous, reproducible evaluations-the scenarios, metrics, and scorecards that gate what ships.
- Leads the architecture of an AI-first, agentic-first evaluation platform in which autonomous agents are first-class operators, designing the services, pipelines, and tooling that agents can operate, extend, and reason about, with the observability and guardrails required to make agent-driven operation trustworthy.
- Owns the platform's hardest challenge: delivering trustworthy results across a deep chain of dependencies pushed far beyond their normal operating envelope, with designs for graceful degradation, intelligent retries, dependency-aware gating, and reliability under sustained and growing load.
- Defines how Copilot's agentic experiences are evaluated end to end, from multi-turn trajectories and tool selection to task outcomes and modalities such as UX and voice, building the simulation, scraping, and scoring capabilities needed to make agent behavior measurable.
- Leads by example and mentors engineers across teams to build extensible, maintainable systems, driving modernization of the evaluation runtime and platform so it can scale with rapidly growing demand while improving cost efficiency and latency.
- Stays at the forefront of agentic systems, large-scale AI evaluation, and autonomous operations, bringing emerging patterns and practices into the organization and sharing that knowledge to raise the bar for how Copilot measures quality.
QualificationsRequired Qualifications:- Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
Other Requirements:Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings:
- Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.
Preferred Qualifications:- Master's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 12+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.
- Deep experience with large-scale distributed systems and reliability engineering, especially delivering dependable outcomes across many interacting service dependencies operating under heavy or unusual load.
- Experience building AI-first or agentic systems, and designing systems, APIs, and tooling intended to be operated by autonomous agents, including the observability, guardrails, and safety mechanisms that make agent-driven operation trustworthy.
- Experience with AI/ML evaluation, measurement, or benchmarking systems, such as LLM-based grading, metric design, or large-scale data and quality pipelines.
Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.