We are seeking a Software Engineer - Scientific Evaluation to own a shared platform for classical testing, scientific benchmarking, and agentic evaluation. The portfolio spans CUDA and C++ libraries, Python packages, PyTorch integrations, scientific models, and AI agents. This hands-on role combines production software engineering, rigorous measurement, distributed systems, and large GPU fleets.
What you will be doing:- Own the architecture and roadmap for evaluation and benchmarking of scientific software, serving agentic and numerical evaluations.
- Design and implement infrastructure, benchmarks and test levels for classical software and scientific agents, balancing rapid turnaround time and thorough scientific coverage.
- Establish trusted references, numerical tolerances, calibrated scorers, regression thresholds, and human-review hooks for nondeterministic workloads.
- Operate heterogeneous GPU capacity using multiple control plane technologies, self-hosted runners, schedulers, queues, containers, caching, observability, and automated recovery.
- We partner with applied scientists, kernel and framework engineers, product teams, and release owners to create reproducible quality signals and production-readiness gates.
What we need to see:- A BS or MS, or equivalent experience, in Computer Science, Computer Engineering, or a related field.
- 5 plus years of relevant industry experience.
- Strong CS fundamentals and production C++ and Python skills, with fluency in Linux, Git, GitHub/GitLab pipeline orchestration, CMake, Docker, Python packaging, and containers.
- A record of creating test, benchmark, evaluation, or distributed execution systems that deliver versioned, reproducible results across repositories.
- Hands-on operation of shared GPU compute with Slurm, Kubernetes, or a similar scheduler, including monitoring, isolation, and failure recovery. Sound measurement practices cover correctness, variance, flakiness, scorer calibration, and regression detection.
Ways to stand out from the crowd:- Background in agentic or LLM evaluation is valuable, especially tool-use tasks, sandboxed execution, trace analysis, model-assisted scoring, and calibration. Multi-GPU or multi-node systems, CUDA software development, statistical benchmarking, and Nsight profiling can also distinguish an application.
- Creative, collaborative engineers who care about reliable science are encouraged to apply!
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until August 13, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
#deeplearning