ServiceNow

Staff Machine Learning Engineer, Agent Eval Platform

ServiceNow$150K — $180K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in applied ML, data science, or related engineering with proven results in production.
  • Experience transforming subjective human judgement into actionable measurements.
  • Strong fundamentals in applied ML and experience evaluating and fine-tuning LLMs.
  • Proficient in Python with a focus on production-grade code rather than exploratory work.
  • Excellent communication skills for articulating complex issues and metrics clearly.
  • High ownership and fast-paced delivery akin to startup environments.
  • Comfortable navigating ambiguity and discerning when measurements are sufficiently reliable.

Responsibilities

  • Design and calibrate a shared base judge with rubrics adjacent to datasets.
  • Define clear problem boundaries for deterministic and LLM-based evaluations.
  • Develop confidence-reporting scoring systems that escalate uncertain judgments.
  • Establish calibration loops with human annotations to enhance judge reliability.
  • Fine-tune small models when necessary to improve evaluation accuracy.
  • Monitor and safeguard against correlated blind spots in LLM evaluations.
  • Analyze and resolve discrepancies between simulated and real-world performance.

Benefits

  • Flexible work environment with options for remote or in-office work.
  • Encouragement of personal ownership and initiative in project delivery.
  • Access to a collaborative team focused on pioneering AI applications.
Full Job Description
Job Description

The Role

Moveworks' AI agents don't just generate text - they act. They plan, call tools, and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did - across a multi-step trajectory through a world it changed - precisely enough that the score can teach it to do better?

That signal is what this role owns. You'll build the judgement layer of our agent evaluation platform: the rubrics, the judges, the calibration against human labels, the methodology that makes a score mean something. And the payoff is larger than a report card - a judge good enough to grade a trajectory is a judge good enough to train against. The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing.

This isn't a pretraining role, and it isn't a testing role. It's applied ML at a point where the methodology genuinely isn't settled: LLMs judging LLMs is an open research problem, and we're working it against agents that take real, irreversible actions in stateful, multi-tenant enterprise environments.

What you get to do in this role:

Judge design and calibration
  • A shared base judge with per-item rubrics expressed as configuration next to the dataset - so eval authors express intent, rather than forking a prompt per eval
  • Splitting the problem correctly: deterministic validators for checkable world state ("was the ticket created, with the right item, routed to the right approver?"), and an LLM judge for the parts that are genuinely fuzzy - was the clarifying question appropriate, was policy followed, was the path efficient
  • Scoring that reports its own confidence, so uncertain judgements route to a human instead of quietly becoming training data
  • A standing calibration loop against human-labeled trajectories, run in partnership with our annotation team - they own the human labeling, you own the calibrated judge artifact. How consistently humans agree with each other sets the ceiling on how good any judge can be, so raising that ceiling is part of the job
  • Fine-tuning a small judge model where an off-the-shelf one isn't good enough
  • Guarding against correlated blind spots: our user simulator and our judge are both LLMs, and they can be wrong in the same direction
  • Offline↔online divergence: when simulation and production disagree, being the person who can say why, and keeping the suite re-seeded from new production failures so it can't quietly overfit

Self-learning for the agent harness

This is where the pillar is headed, and a large part of why the seat exists.
  • A calibrated trajectory judge is, functionally, a reward model. Turning ours into a process reward model - a dense, step-level signal for what a good agent trajectory looks like - is the unlock
  • Using that signal to optimize the agent itself: prompts, tool selection, planner behavior, retrieval, routing - tuned against simulation rather than against production traffic
  • Building the substrate a future RL effort runs on: versioned scenarios, a repeatable simulated world, and a reward signal calibrated to human judgement
  • Holding the guardrail that keeps this honest: step-level scores train and diagnose; end-state outcomes are what we hold the agent to. Scoring individual steps is powerful for attribution and as a training signal, and dangerously brittle as a definition of success


Qualifications

To be successful in this role you have:
  • 8+ years in applied ML, data science, or ML-adjacent engineering, with a track record of work that shipped and got used
  • Experience turning subjective human judgement into a measurement that holds up - one that other people, and ideally other models, can act on. This is the core of the job
  • Strong applied ML fundamentals, and comfort treating LLMs as a component you evaluate, prompt, and fine-tune rather than one you pretrain
  • Strong Python, and the discipline to ship production-grade code rather than notebooks
  • Ability to think and communicate clearly about complex problems - a large part of this job is convincing engineers that a number means what you say it means, and being right
  • A high degree of ownership and a bias toward shipping at startup pace
  • Comfort with ambiguity, and the judgement to know when a measurement is good enough to act on

Experience in at least 3 of these:
  • LLM-as-judge or automated evaluation design, and calibrating it against human judgement
  • Human annotation programs: rubric authoring, label quality, and annotator throughput as a real constraint
  • Search ranking, recsys, or online experimentation evaluation - golden-set staleness, offline/online divergence, side-by-side rater agreement. This is the closest existing analog to agentic eval, and it transfers directly
  • Fine-tuning and evaluating small models: SFT, preference tuning, distillation
  • Reward modeling, RLHF/RLAIF, or process reward models
  • Agent trajectory analysis and step-level fault attribution
  • Prompt engineering as an engineering discipline - versioned, tested, and measured, not tuned by vibes


Additional Information

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

About ServiceNow

ServiceNow provides cloud-based solutions that define, structure, manage, and automate services for enterprise operations in North America, Europe, the Middle East, Africa, the Asia Pacific, and other countries. The company offers service management solutions, including incident, problem, change, request, and cost management as well as service catalogs; and IT, HR, facilities, and field service management solutions. It also provides IT operations management solutions covering service mapping, delivery, and assurance solutions; business management solutions such as financial management, project portfolio suite, vendor performance management, and performance analytics as well as governance, risk, and compliance; and application development services.

ServiceNow Careers

Join the dynamic team at ServiceNow, a global leader in digital workflow solutions, where innovation and leadership converge to shape the future of work. At ServiceNow, we offer more than just job opportunities; we provide a platform for professional growth and a chance to be part of a culture that values diversity, creativity, and continuous learning.

Work You’ll Do

Embark on a career journey with ServiceNow and contribute to the world’s leading enterprises' digital transformation. Our team is at the forefront of developing cutting-edge technologies that improve how people work. With ServiceNow, you will use your skills to impact businesses and industries profoundly, driving efficiency and innovation.

Join Our Market-Leading Team

ServiceNow is not just another technology company. We are a team that thrives on diversity and leadership, fostering an inclusive environment that promotes growth and development. Our commitment to diversity training ensures that every team member can achieve their potential.

Innovative Work

ServiceNow is home to more than 10,000 dedicated professionals who lead the charge in digital workflows and enterprise solutions. As part of our team, you will engage in projects that merge technology with practical applications, creating revolutionary products that advance how services are delivered and managed.

Career Development

At ServiceNow, your career trajectory is filled with boundless opportunities. We support your growth with robust training programs, leadership development courses, and access to global challenges. Whether you are looking for an internship, full-time position, or leadership role, ServiceNow equips you with the tools to excel.

Be Part of a Great Team

Working at ServiceNow means being part of a community that values teamwork and innovation. Our collaborative environment encourages networking and sharing ideas, making our workplace vibrant and dynamic. The benefits of joining ServiceNow extend beyond comprehensive health and wellness; they include fostering professional connections and friendships that last a lifetime.

Explore Job Opportunities and Internships

Whether you’re a seasoned professional or a recent graduate, ServiceNow offers a range of employment options to suit your career goals. From internships that provide real-world experience to full-time positions that challenge you to leverage your expertise, we are committed to hiring the best talent.

Stay Connected

Join Our Team Search open positions that match your skills and interests. At ServiceNow, we look for passionate, curious, and solution-driven team players. Explore the possibilities that await you at a company that is committed to your professional success.

Keep Up to Date

Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here.

Job Alert Emails

Customize your subscription to receive job alerts, the latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities that await at ServiceNow.

ServiceNow Careers

Empowering professionals to achieve more, ServiceNow is where careers are future-proofed, and ambitions are realized. Join us in our journey of growth and innovation.
Learn more about ServiceNow
Size
16,881 employees
Market Cap
$76.5 billion
Industry
Net Income
$118.5 million
Founded
2004
5 Year Trend
+33.5%
Revenue
$4.5 billion
NASDAQ

Similar Jobs

More Jobs at ServiceNow

More Information Technology Jobs

Find similar Staff Machine Learning Engineer, Agent Eval Platform jobs: