Leonardo DRS

Applied AI Lead - Evaluation & Measurement

Leonardo DRS$120K — $145K *
Aerospace & Defense
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in data, ML, or test engineering with 3+ years in evaluation or model validation.
  • Experience halting a project due to evaluation concerns, as well as advocating for a project's launch.
  • Established a reliable measurement practice with a focus on operational continuity.
  • Proficient in data instrumentation, including tracking pipelines and creating dashboards.
  • Proven record of professional integrity, especially under pressure with senior stakeholders.

Responsibilities

  • Define evaluation criteria for AI capabilities, including error budgets and documentation before launch.
  • Build measurement pipelines to track AI capability performance.
  • Establish connections between measurements and the source systems for accuracy.
  • Set standards for logging and telemetry in deployed systems for audits.
  • Deliver accurate monthly and quarterly performance reports to finance without discrepancies.
  • Manage the intake of AI ideas to ensure all proposals are officially captured and handled.
  • Create tools to validate claims while remaining independent from those making the claims.

Benefits

  • Comprehensive medical, dental, and vision coverage.
  • Company-contributed health savings account.
  • Telemedicine options and life/disability insurance.
  • 401(k) savings plan with employer contributions.
  • Wellness programs focused on physical, emotional, and financial well-being.
  • Flexible work schedules and competitive vacation offerings.
Full Job Description
Job ID: 115203

This is a hybrid position open to candidates who reside in or near Beavercreek, Ohio; Frederick, Maryland; or Fort Walton Beach, Florida.

Job Summary

Prove AI works. Prove what it's worth.

We've used AI before, now we treat it as a first-class citizen at Airborne & Intelligence Systems (AIS). We build and deploy AI across signals intelligence, electronic warfare, air combat training, and the mission systems that connect them. This work shows up on our floor, in the field, and on the company's P&L, while meeting the demanding standards of our defense and commercial customers. If you want your work and your talent to count - this is the place.

Two questions follow every AI system we build: does it actually work, reliably on real jobs and is it actually worth the money we say it is? You own both answers. You design the test every capability must pass before it can ship, and you build the measurement that turns results into honest numbers leadership can trust. You hold the line. Nothing ships until it clears your bar - no test, no launch - and you're deliberately independent from the person who builds and deploys, so the check is real. When there's pressure to over-claim a result, you're the one who holds the number honest.

At the heart of this function is a proprietary platform, the encoded way this business does AI. It takes any employee from a rough idea to a governed, well-formed use case, checks it against the rules, and prepares a recommendation a human decision-maker rules on. You don't just use it; you also build it. The quality bar and measurement logic you design get embedded into the platform itself, so the standard travels with the system to every team that touches it. You encode the acceptance standard and the value math once, and every team inherits both.

This is a high impact, founding individual-contributor seat within the technical core of AI Excellence function with meaningful end to end work ownership. You report directly to top leadership of the function giving you outsize visibility, opportunity for personal growth and direct line to impact. Your peers will be AI leads for Solutions & Delivery/ Infrastructure & Deployment

Job Responsibilities

Primary & Essential Accountabilities

  • Define "good enough to ship." Design the evaluation for each capability - baseline, acceptance criteria, error budget, human-review points - documented before launch. Gate the launch.
  • Build the measurement. Stand up the pipelines and record keeping that track what each capability actually delivers.
  • Establish the truth. Connect each measurement directly to the source system that holds the ground truth.
  • Define what deployed systems must log. Set the trace and telemetry standard that keeps every shipped capability diagnosable and auditable.
  • Produce the numbers. Deliver the monthly and quarterly results - forecast versus actual, clean and without double-counting - that the function defends to Finance.
  • Keep the front door open. Run the record intake so every AI idea is captured, and nothing happens off the books.
  • Scope boundaries. You are not the claim owner, governance or reporting. You build the instrument that proves or disproves the claims in an auditable manner.


Qualifications

  • 5+ years in data, ML, or test engineering (regardless of total years exp), including 3+ years in evaluation or model validation, with recent hands-on large-language-model evaluation work.
  • Demonstrated bi-directional evaluation rigor: an example where your evaluation stopped something from shipping, and one where it cleared something others doubted. Expect us to probe the methodology of each.
  • Built a measurement practice where none existed. The first baselines, the first pipelines, the first dashboard leadership actually used, and it kept producing after you handed it off.
  • Data and measurement instrumentation skills: pipelines, dashboards, record-keeping, baselines from source-of-record data.
  • A demonstrated record of professional integrity under pressure: you have told someone, especially seniors "the number doesn't support that claim" and held.
  • Bachelor's degree and 5 years experience or equivalent combination of both. Engineering, Computer Science, Data Science, or a related technical field is preferred
  • Five years of relevant experience (or an equivalent combination of experience and training that provides the required knowledge, skills, and abilities
  • Proficient technical expertise with demonstrated application
  • Excellent interpersonal, leadership, negotiation, communication and writing skills


Preferred Qualifications

  • Experience wiring metrics to operational systems of record (not survey- or estimate-based measurement).
  • Statistical literacy for probabilistic acceptance: error budgets, confidence bounds, claims that survive challenge.
  • Finance-adjacent exposure: cost accounting, benefits realization, or audit (a plus, not a requirement).
  • Defense cost-accounting context.
  • Active or prior security clearance.


U.S. Citizenship required. This position requires an active DOD security clearance or the ability to obtain such clearance within a reasonable time after commencement of employment.

Taking care of our people is a top priority at Leonardo DRS. We are proud to offer competitive salaries and comprehensive benefits, including medical, dental, and vision coverage, a company contribution to a health savings account, telemedicine, life and disability insurance, legal insurance, and a 401(k) savings plan. We champion wellness programs that focus on physical, emotional, and financial well-being. We develop our talent by offering programs and activities to support career-growth, professional development, and skill enhancement. And we understand there is more to life than work, and the importance of offering flexible work schedules with our 9/80 program, competitive vacation, health/emergency leave, paid parental leave, and community service hours.
*Some employees are eligible for limited benefits only

About Leonardo DRS

Leonardo DRS is a leading supplier of integrated products, services and support to military forces, intelligence agencies and prime contractors worldwide. Focused on defense technology, we develop, manufacture and support a broad range of systems for mission critical and military sustainment requirements, as well as homeland security. Headquartered in Arlington, Virginia, the Company is a wholly owned subsidiary of Leonardo S.p.A. which employs more than 45,000 people worldwide. We offer a competitive compensation package and a business culture that rewards performance. For additional information on Leonardo DRS, please visit our website at www.leonardodrs.com.
Learn more about Leonardo DRS
Size
11,000 employees
Market Cap
$3.1 billion
Industry
Founded
1968
5 Year Trend
+55.7%
NASDAQ

Similar Jobs

More Jobs at Leonardo DRS

More Aerospace & Defense Jobs

Find similar Applied AI Lead - Evaluation & Measurement jobs: