Engagement Manager, Agentic AI Workflow Evaluations

Innodata, Inc.

$156K — $177K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree or equivalent practical experience
  • 5+ years of experience managing a delivery team of 10-20 people
  • Proven record as the customer contact for a services engagement with enterprise or tech clients
  • Extensive background in AI/ML evaluation tasks such as red-teaming or trust and safety review
  • Ability to critically evaluate rubric-scored work and provide independent assessments
  • Strong written communication skills for drafting reports and escalation memos

Responsibilities

  • Own end-to-end delivery for the onsite evaluation team
  • Manage recruitment, onboarding, and performance of team members
  • Act as the primary point of contact for the client's program leads
  • Audit team outputs and ensure scoring justifications meet client expectations
  • Collaborate with QA lead on quality metrics and rubric evaluations
  • Oversee escalation of safety and ambiguous findings to the client
  • Identify potential growth opportunities within the engagement

Benefits

  • Opportunity to develop leadership skills in a cutting-edge AI environment
  • Engagement with a diverse team focused on complex AI evaluations
  • Exposure to high-level enterprise clients and their technical leads
  • Chance to influence the growth and direction of projects
  • Commitment to maintaining information security and privacy standards
Full Job Description
Scope of the Role:

We are standing up a dedicated onsite team to evaluate complex, real-world agentic AI workflows for a frontier AI customer. The Engagement Manager owns this team and this engagement: a group of ten reviewers and one QA lead working inside isolated test environments, assessing whether AI agents complete tasks safely, respect user intent and consent, and hold up under close scrutiny.

This is a delivery leadership role with a strong customer-facing component. You are the customer's primary day-to-day contact at the site, accountable for throughput, quality, and turnaround, and responsible for identifying where the engagement should grow next. You will not be scoring trajectories as a daily task, but you need enough technical depth to audit the team's work yourself and defend a scoring decision in a room with the customer's technical leads.

What You'll Own:
  • Own end-to-end delivery for the onsite team: throughput, quality, turnaround, and capacity against committed volumes
  • Manage and develop a team of eleven, including hiring, onboarding, performance management, and retention
  • Serve as the primary onsite point of contact for the customer's program and technical leads; run recurring reviews, report on delivery metrics, and resolve issues before they escalate
  • Audit reviewer output directly - sample trajectories, check rubric application, and assess whether scoring rationale would survive customer review
  • Partner with the QA lead on calibration cycles, quality metrics and evaluation, and rubric refinement; arbitrate disagreements that calibration does not resolve
  • Own the escalation path for safety-relevant and ambiguous findings, including judgment on what reaches the customer and how quickly
  • Identify expansion opportunities within the account: adjacent workstreams, new task types, additional capacity, and scope the work with internal delivery and commercial teams
  • Forecast staffing and cost against the engagement's commercial model, and flag variance early
  • Maintain information security, privacy, and facility access practices required by the customer's onsite environment

You'll Thrive in This Role If You Have:
  • Bachelor's degree or equivalent practical experience
  • 5+ years of professional experience, including direct management of a delivery team of 10-20 people
  • Track record as the customer-facing owner of a services or delivery engagement with an enterprise or technology client, including metrics reporting and issue escalation
  • Significant experience with AI/ML evaluation work: annotation, red-teaming, RLHF, model or agent evaluation, or trust and safety review
  • Ability to read rubric-scored work critically and form an independent view of whether a score is correct
  • Strong written communication; able to produce reporting and escalation memos that a technical customer will accept without rework

The expected hourly salary range for this position is $75-85 p/hour, based on experience, skills, and qualifications.

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission's guide at https://consumer.ftc.gov/articles/job-scams.

If you believe you've been targeted by a recruitment scam, please report it to Innodata at [redacted] and consider reporting it to the FTC at ReportFraud.ftc.gov.

Similar Jobs

More Jobs at Innodata, Inc.

More Information Technology Jobs

Find similar Engagement Manager, Agentic AI Workflow Evaluations jobs: