Product Manager, Evals & Improvement

Meta

$150K — $180K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Strong quantitative skills with experience in defining complex metrics.
  • Ability to navigate ambiguous problem spaces lacking established metrics.
  • 5+ years in the industry, including 2 years in Product Management.
  • Bachelor's degree in a STEM field preferred but not essential.
  • Experience working closely with data science and ML teams.
  • Proven track record of using data and experimentation for product improvements.
  • Direct experience in building end-to-end evaluation frameworks for AI.

Responsibilities

  • Own the full lifecycle of feedback and improvement processes.
  • Define measurement criteria for agent quality across various dimensions.
  • Develop mechanisms to capture user signals at scale.
  • Collaborate with ML and data science teams to improve models based on feedback.
  • Establish evaluation strategies for metrics of success and regression detection.
  • Ensure transparency and accountability for quality within the product team.

Benefits

  • Opportunity to shape the future of work through innovative AI solutions.
  • Work in a dynamic, high-impact role that directly influences revenue and growth.
  • Collaborate cross-functionally with engineering and data science teams.
  • Access to cutting-edge resources and tools in AI technology and product development.
  • Continuous learning and professional development opportunities in a fast-evolving field.
Full Job Description
The Agent Transformation Accelerator (ATA) Product team is building Meta's future of work: Metamate, an internal full-stack platform that lets people direct, review, and help agents improve their work across the entire company. We're building the platform components, memory systems, improvement loops, and shared product experiences that make agentic AI genuinely useful for product development. This role owns how we measure and raise the quality of agentic systems (multi-step trajectories, tool use, and partial credit). This is an early-stage, high-leverage role that contributes directly to Metamate's topline revenue, growth, and velocity. You'll shape the approach from the ground up in a space where the right metrics don't yet exist, defining what "good" means for agents and owning the eval verdict that gates model upgrades, harness changes, and major launches. Expect to be constantly learning, working shoulder-to-shoulder with engineering, data science, and ML, including our partners across ATA and Meta Superintelligence Labs (MSL), to turn signal from real usage into a continuous improvement loop that makes the product measurably better every week.

Responsibilities

Own the end-to-end feedback and improvement system: collection, analysis, prioritization, and resolution.
• Define how we measure agent quality across dimensions (accuracy, helpfulness, reliability, safety).
• Build product mechanisms that capture implicit and explicit user signal at scale.
• Partner with ML and data science to turn feedback into model improvements, eval frameworks, and training data.
• Drive the evals and measurement strategy: what "good" looks like, how we detect regressions, and how we track progress.
• Create transparency and accountability for quality across the ATA Product team.

Minimum Qualifications
• Strong quantitative skills and experience defining complex metrics
• Comfort with ambiguous problem spaces where the right metrics do not yet exist
• 5+ years of relevant industry experience with at least 2 years in Product Management
• Bachelor's degree (or relevant degree equivalent): STEM subject ideal but not essential (Computer Science, Engineering, Information Systems, Analytics, Mathematics, Physics, Applied Sciences)
• Experience partnering closely with data science and ML engineering teams
• Track record of driving product improvements through data and experimentation
• Direct experience building evals for AI products end to end: task set construction, rubric design, instrumentation, and analysis
• Experience running eval-driven development cycles, with a clear view of where evals are informative and where they mislead
• Willingness to get into the weeds of eval data and rubrics, with high standards for eval rigor and an obsession with quality
• Deep understanding of LLM capabilities and failure modes

Preferred Qualifications
• Experience building annotation, labeling, or crowd-sourcing systems
• Experience with reinforcement learning from human feedback (RLHF) or similar human-in-the-loop systems
• Background in evals, trust and safety measurement, or ML quality infrastructure
• Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
• Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
• Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies

Similar Jobs

More Jobs at Meta

More Enterprise Technology Jobs

Find similar Product Manager, Evals & Improvement jobs: