We are looking for an Evaluation/ML-Systems Engineer to own how we measure the program. You will play a critical role in separating real capability from anecdote. You will make our numbers mean something, and keep them meaning it as the program grows. Your benchmarks will define what progress means for the program. Every conclusion the program reaches should trace back to benchmarks, runs, and code you helped make reproducible.
Measurement is not a support function here; it is how the program knows what is true. You will design the metrics, choose the baselines, and write the protocols that make comparisons fair. You will build the infrastructure that keeps every result tied to the run that produced it. When a number looks surprising, your systems will make it easy to check. You will automate the boring parts of measurement so people can focus on the questions.
You will partner with security researchers and platform engineers to design experiments people can trust and rerun. We review results internally before we rely on them, and your work makes that review possible. You will also help the team learn to read its own evidence honestly. Together you will build a shared habit of checking claims before repeating them.
What You'll Be Doing:- Evaluation infrastructure: Build the benchmarking and reproducibility systems we depend on.
- Metrics and protocols: Define the metrics and protocols we measure against.
- Traceability: Map every result to the code and runs that produced it.
- Evidence discipline: Keep findings reviewable and conclusions traceable.
What We Need To See:- Bachelor's degree (or equivalent experience) with 5+ years in ML engineering or evaluation.
- Evaluation experience: Designing benchmarks, metrics, and statistically sound comparisons for ML systems.
- Measurement rigor: A careful, skeptical approach to metrics, baselines, and claims.
- Engineering skills: Solid Python engineering for shared infrastructure, including experiment tracking and data pipelines.
Ways to Stand Out from the Crowd:- Security evaluation: Exposure to evaluating security tooling or pipelines.
- Agentic systems: Experience measuring agent or LLM behavior.
- Community work: Contributions to public benchmarks or evaluation frameworks.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until July 30, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.