Wipro

AI Evaluation Engineer

Wipro • $60K — $148K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Strong software engineering background with experience in automation and developer tooling.
  • Experience in evaluating AI coding agents and validating generated code changes.
  • Skilled in building reproducible evaluation workflows, including test execution and result validation.
  • Proficient in benchmarking and reliability measurement, understanding run-to-run variance and failure analysis.
  • Expertise in Git, CI/CD integration, and containerization using Docker.
  • Familiarity with using LLMs to evaluate coding-agent outputs and calibration against human assessments.
  • Hands-on experience with AI coding tools like Claude Code, Devin, or Cursor and their best practices.

Responsibilities

  • Build and integrate evaluation harnesses for software development use cases.
  • Create versioned, repeatable processes to evaluate AI tools and ensure environmental consistency.
  • Validate evaluation methods against human judgment for consistent scoring accuracy.
  • Conduct execution-based benchmarking across quality, productivity, and efficiency measures.
  • Analyze performance data across repeated runs, focusing on variance and reliability of workflows.
  • Collaborate with engineering and data teams to enhance tooling effectiveness and document methodologies.

Benefits

  • Full range of medical and dental benefits options.
  • Disability insurance.
  • Paid time off including sick leave.
  • Additional paid and unpaid leave options.
Full Job Description
Job Title: AI Evaluation Engineer

City: San Diego

State/Province: California

Posting Start Date: 9/22/26

Job Description:

Job Description

Role Overview

Help build and scale the tooling we use to measure how well AI-powered software development tools actually perform. You'll develop evaluation harnesses, automate benchmark runs, and help make sure the results we produce are reproducible and hold up to scrutiny. This is an engineering role, but a lot of the work is about getting the measurement right, not just automating it.

͏

Key Responsibilities
• Build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.
• Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated worktrees) so results stay comparable over time.
• Validate and calibrate evaluation approaches against human judgment, so scores are consistent and correct rather than just repeatable.
• Support execution-based benchmarking across quality, productivity, and e iciency measures, including cost and latency.
• Analyze results across repeated runs, looking at variance, failure patterns, and cost per outcome, and find ways to make the workflows more reliable and more automated.
• Work with engineering and data teams to improve the tooling, and document how the evaluations work and what they found for both technical and leadership audiences.

͏

Required Skills & Experience
• Strong software engineering background, with real experience building automation, developer tooling, or test and validation systems.
• Coding-agent evaluation: Evaluating AI coding agents that modify code repositories, including validating generated code changes against expected outcomes.
• Evaluation harnesses: Building automated, reproducible evaluation workflows, including test execution, environment setup, and result validation.
• Benchmarking and reliability: Establishing baselines, measuring run-to-run variance, analyzing failures, and ensuring consistent evaluation results.
• Git and CI/CD integration: Working with repository history, branches, pull requests, and automated testing within CI/CD pipelines.
• LLM-as-a-Judge: Using LLMs to evaluate coding-agent outputs, including calibration against human assessments.
• Proficient in at least one general-purpose language such as Python, Java, or JavaScript - the specific language background is flexible.
• Solid working knowledge of Git, including how branches, history, and working trees behave, and of containerization with Docker.
• Experience with APIs, development environments, CI/CD pipelines, and typical engineering workflows.
• Understanding of how AI, LLM, or agent evaluation works and where it goes wrong, such as why a judge can be consistent but still wrong, why a single run can mislead, and how benchmark contamination happens.
• Able to troubleshoot technical problems, think clearly about whether a measurement is valid, and analyze results carefully.
• Hands-on experience using AI coding tools and agentic harnesses such as Claude Code, Devin, or Cursor, and command of the best practices for working with them effectively.

Mandatory Skills: Cloud Product & Platform Testing.

Experience: 5-8 Years.

The expected compensation for this role ranges from $60,000 to $148,500 .

Final compensation will depend on various factors, including your geographical location, minimum wage obligations, skills, and relevant experience. Based on the position, the role is also eligible for Wipro's standard benefits including a full range of medical and dental benefits options, disability insurance, paid time off (inclusive of sick leave), other paid and unpaid leave options.

Applicants are advised that employment in some roles may be conditioned on successful completion of a post-offer drug screening, subject to applicable state law.

About Wipro

Wipro Limited is an Indian multinational corporation that provides information technology, consulting and business process services. The company was founded in 1945 and is headquartered in Bengaluru, India. Wipro has operations in over 50 countries and employs over 191,000 people. The company's primary business is in the information technology sector, and it provides services such as application development and maintenance, digital strategy consulting, and data analytics.
Learn more about Wipro
Size
240,000 employees
Market Cap
$25.9 billion
Industry
Net Income
$101.4 billion
Founded
1945
5 Year Trend
+7.5%
Revenue
$614 billion
NASDAQ

Similar Jobs

More Jobs at Wipro

  • Wipro
    AI Evaluation Engineer
    $60K — $148K *
    San Diego, CA 92154 (San Diego County)
    Information Technology
    In-Person
  • Wipro
    Senior project Engineer
    $80K — $160K *
    Vancouver, WA 98682 (Clark County)
    Manufacturing & Automotive
    In-Person
  • Wipro
    Quick Turn Program Manager
    $80K — $160K *
    Santa Clara, CA 95051 (Santa Clara County)
    Manufacturing & Automotive
    In-Person
  • Wipro
    Quick Turn Program Manager
    $80K — $160K *
    Santa Clara, UT 84765 (Washington County)
    Manufacturing & Automotive
    In-Person
  • Wipro
    CREO DEVELOPER L3
    $60K — $140K *
    Sunnyvale, CA 94087 (Santa Clara County)
    Technical Services
    In-Person

More Information Technology Jobs

Find similar AI Evaluation Engineer jobs: