NVIDIA Corporation

Senior Software Engineer, Quantized Inference

NVIDIA Corporation$152K — $287K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Proficient in Python; familiarity with C++
  • Strong software engineering fundamentals: concise, well-tested code; fluent with AI-assisted tooling
  • Experience with ML accelerators; basic understanding of how certain ML layers affect execution time
  • Familiarity with PyTorch internals or equivalent framework
  • Experience with large open-source codebases
  • MS/PhD in Computer Science or related field, or equivalent experience
  • 4+ years in a relevant software engineering role
  • Strong communication skills and ability to navigate ambiguous requirements.

Responsibilities

  • Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang)
  • Own model export pipelines ensuring quantized checkpoints serialize correctly
  • Build prototypes and benchmarking harnesses to evaluate throughput/interactivity
  • Develop data analysis tooling and visualizations for debugging
  • Improve developer productivity across team systems and infrastructure
  • Participate in code reviews and incorporate feedback.

Benefits

  • Eligibility for equity
  • Participation in employee benefits programs.
  • Use of AI tools in recruiting processes
Full Job Description
We are now looking for a Senior Software Engineer for Quantized Inference! NVIDIA is seeking software engineers to accelerate the discovery and deployment of efficient inference recipes for LLMs. A recipe defines which operators are transformed into low-precision or sparsified variants - unlocking throughput and latency gains without regressing accuracy or verbosity. Recipes may incorporate techniques such as rotations, block scaling to attenuate outlier impact, or improved calibration data drawn from SFT/RL pipelines.

Each new recipe demands corresponding kernel and model-level implementations in inference engines (vLLM, TRT-LLM, SGLang). The candidate will translate recipe specifications into functionally correct, performant code, e.g., writing Triton kernels, inserting quantize/dequantize nodes into prefill and decode paths, and ensuring per-expert scaling in MoE layers is handled correctly. From there, the candidate will collaborate with partner inference teams to further optimize throughput and interactivity on target workloads. This work is a core component of our productization effort across Megatron-LM, ModelOpt, and vLLM.

What you'll be doing:
  • Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang)
  • Own model export pipelines (ModelOpt, Megatron-LM <-> HuggingFace), ensuring quantized checkpoints serialize correctly for downstream serving
  • Build prototypes and benchmarking harnesses to evaluate recipe throughput/interactivity before full optimization
  • Develop data analysis tooling and visualizations for numerics debugging
  • Improve developer productivity across the team: CI, build systems, training infrastructure, pipeline friction
  • Participate in code reviews and incorporate feedback


What we need to see:
  • Proficient in Python; familiarity with C++
  • Strong software engineering fundamentals: concise, well-tested code; fluent with AI-assisted tooling
  • Experience with ML accelerators with a basic understanding of how certain ML layers affect execution time
  • Familiarity with PyTorch internals (custom ops, autograd, export) or equivalent framework
  • Experience reading, modifying, or contributing to a large open-source codebase
  • MS/PhD in Computer Science or related field, or equivalent experience.
  • 4+ years in a relevant software engineering role
  • Demonstrated ability to move fast with ambiguous requirements, with strong written and verbal communication


Ways to stand out from the crowd:
  • Experience contributing to inference serving frameworks (vLLM, TRT-LLM, SGLang) or Triton kernel development
  • Track record of debugging numerical issues across mixed-precision boundaries
  • Deep experience with model compression techniques: PTQ, QAT, structured/unstructured sparsity


Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until July 26, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Information Technology Jobs

Find similar Senior Software Engineer, Quantized Inference jobs: