NVIDIA Corporation

Senior Software Engineer - AI Inference Performance

NVIDIA Corporation$184K — $356K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years in full-stack LLM/VLM inference performance with measurable production gains.
  • Strong programming skills in Python, Rust, and/or C++, plus CUDA or GPU programming experience.
  • Expertise in speed-of-light analysis and roofline models using NVIDIA Nsight tools.
  • Deep understanding of GPU architecture, including Tensor Cores and memory hierarchies.
  • Practical experience optimizing inference servers for various workloads.
  • Knowledge of distributed systems and networking for accelerated computing.
  • BS or MS in Computer Science, Computer Engineering, or relevant field.

Responsibilities

  • Lead analysis of LLM/VLM inference processes and define workloads.
  • Optimize performance metrics such as latency and efficiency across models.
  • Build performance models to assess optimization opportunities.
  • Profile workloads and eliminate bottlenecks in various code aspects.
  • Tune hyperparameters and optimization techniques relevant to workload.
  • Develop and optimize performance-critical kernels using CUDA technologies.
  • Establish benchmarks and manage performance regression metrics.

Benefits

  • Eligible for equity participation.
  • Access to high-impact AI projects and contributions.
  • Collaborative work with cutting-edge technologies.
Full Job Description

You will push workloads toward practical performance limits on NVIDIA GPU-accelerated systems. Your work will span models, serving software, distributed runtimes, communication, CUDA kernels, and GPU architecture. Deliver measurable gains in latency, throughput, efficiency, and scale.

This is a hands-on role for an engineer who turns performance models and profiler data into working code. You will collaborate with model, framework, kernel, networking, and GPU architecture teams. You will contribute improvements to open-source inference engines and develop methods that others can reproduce. Your work will improve production deployments and help build future NVIDIA platforms.

What you'll be doing:

  • Lead end-to-end analysis of LLM/VLM inference processes. Define representative prefill and decode workloads. Optimize time to first token, inter-token latency, P99 end-to-end latency, processing efficiency, and key-value (KV) cache capacity. For multimodal models, isolate preprocessing, encoder, and decoder costs.

  • Build speed-of-light and roofline models to quantify performance headroom. Connect arithmetic intensity, bandwidth, occupancy, memory hierarchy, and communication costs to clear optimization hypotheses.

  • Profile workloads using NVIDIA Nsight Systems, Nsight Compute, PyTorch Profiler, and custom instrumentation. Eliminate bottlenecks in host code, CUDA kernels, memory, communication, and scheduling.

  • Tune serving hyperparameters and techniques such as batching, KV-cache management, quantization, speculative decoding, CUDA Graphs, and model parallelism. Choose them based on workload, hardware, model quality, and service-level objectives.

  • Build and optimize performance-critical kernels, including attention, matrix multiplication, mixture-of-experts routing, quantization, and data movement. Use CUDA, CUTLASS, Triton, or related technologies.

  • Establish repeatable benchmarks, canonical run records, and performance regression gates. Manage aspects such as model, precision, hardware, topology, software, features, and workload; Balance between performance and accuracy. Collaborate across with various teams and contribute high-quality upgrades to TensorRT-LLM, vLLM, SGLang, or associated projects.

What we need to see:

  • More than 6 years of experience in full-stack LLM/VLM inference performance involving models, serving, distributed runtimes, kernels, and hardware. Your efforts result in measurable gains in production or production-representative environments.

  • Strong programming skills in Python, Rust and/or C++, plus hands-on experience with CUDA or another GPU programming environment.

  • Demonstrated expertise in speed-of-light analysis, roofline models, microbenchmarks, and tools including NVIDIA Nsight Systems and Nsight Compute. You convert profiles into testable hypotheses and validated progress.

  • Deep understanding of GPU architecture, including Tensor Cores, memory hierarchy, caches, occupancy, synchronization, and numerical formats across hardware generations.

  • Practical experience optimizing inference servers and model execution. You can choose techniques for the workload, including batching, scheduling, KV-cache management, quantization, speculative decoding, and various parallelism strategies

  • Understanding of distributed systems and networking for accelerated computing. You can reason about collectives, topology, and scale-up versus scale-out performance.

  • BS or MS in Computer Science, Computer Engineering, or a related field, or equivalent experience.

Ways to stand out from the crowd:

  • Contributions to one or more high-performance AI projects. Examples include TensorRT-LLM, vLLM, SGLang, PyTorch, CUDA, Triton, or NCCL.

  • Experience developing AI-agent-supported performance workflows that automatically gather and analyze profiles, identify bottlenecks, explore serving configurations, or produce optimized runtime and kernel code. You validate generated changes through reproducible, human-reviewed tests for performance, model quality, and correctness.

  • Published research, conference presentations, technical talks, or blog posts that clearly explain inference performance methods and results.

  • Delivered advancements for new LLM or VLM architectures, long-context inference, mixture-of-experts models, multimodal pipelines, or large-scale distributed serving.Successfully carrying these out will shape the future of AI inference performance!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and .

Applications for this job will be accepted at least until August 30, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Information Technology Jobs

Find similar Senior Software Engineer - AI Inference Performance jobs: