NVIDIA Corporation

Senior Software Engineer, AI Inference

NVIDIA Corporation$135K — $220K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or equivalent experience
  • 5+ years of experience with complex, production-grade software systems
  • Hands-on experience deploying LLM inference workloads, specifically with vLLM
  • Proficient in container orchestration (Kubernetes) and HPC scheduling (Slurm)
  • Strong understanding of LLM serving fundamentals, including batching strategies and KV cache management
  • Experience with GPU performance analysis tools like Nsight Systems and Nsight Compute
  • Excellent communication skills to convey technical findings to various audiences

Responsibilities

  • Engage with customer engineering teams to understand LLM serving architectures and performance goals
  • Design and implement benchmarking campaigns for Kubernetes and Slurm environments
  • Optimize vLLM deployments on GPU clusters for performance
  • Develop performance improvement plans based on profiling insights
  • Create tools and automation to enhance team and customer productivity
  • Document technical findings and recommend improvements for open-source projects

Benefits

  • Comprehensive benefits package for employees and their families
  • Access to professional development resources and tools
  • Flexible work arrangements support including hybrid options
  • Equity options as part of compensation
  • Opportunities to work on impactful open-source projects
Full Job Description
Senior Software Engineer, AI Inference Help us push the boundaries of AI inference at NVIDIA - where your systems expertise shapes both the technology and the teams building on top of it! We're looking for a Senior Software Engineer to work at the frontier of large-scale LLM serving, partnering directly with some of the world's most technically demanding customers to unlock the full performance potential of NVIDIA's inference stack. In this role, you'll combine deep systems knowledge with hands-on customer engagement - profiling real deployments, benchmarking across GPU clusters, and turning insights into improvements that ripple across the open-source ecosystem. Do you love digging into performance problems that don't have obvious answers, and want your work to have an impact far beyond a single codebase? We'd love to talk. Unlike traditional customer-facing engineering roles, we expect you to go far deeper - contributing to vLLM, NVIDIA Dynamo, and the tooling that makes every engineer on your team more effective. What You'll be doing: • Work directly with customer engineering teams through long-term technical partnerships, understanding their LLM serving architectures and performance goals, then designing and implementing end-to-end benchmarking campaigns across Kubernetes and Slurm environments to surface actionable insights. • Set up and operate vLLM serving deployments on GPU clusters, tuning configurations for throughput, latency, and efficiency - and collect Nsight Systems / Nsight Compute profiling traces to identify performance gaps relative to reference frameworks. • Develop detailed performance plans based on profiling findings and collaborate with NVIDIA's kernel engineering and OSS vLLM teams to drive improvements that benefit both your customers and the broader community. • Build internal tools, benchmarking harnesses, and automation pipelines that raise the productivity of your teammates and customers alike - with a multiplier attitude that makes everyone around you more effective. • Document architectures, findings, and recommendations with clarity for technical audiences, and contribute improvements back to vLLM and related open-source projects where appropriate. What We Need to See: • Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, or equivalent experience. • 5+ years of industry experience building and operating complex, production-grade software systems, with strong instincts for how systems behave at scale. • Hands-on experience deploying and operating LLM inference workloads - particularly with vLLM - including configuration, optimization, and debugging in real-world environments. • Proficiency with container orchestration (Kubernetes) and HPC scheduling (Slurm) for running GPU-accelerated workloads. • Solid understanding of LLM serving fundamentals: batching strategies (continuous batching, chunked prefill), KV cache management, and tensor/pipeline parallelism. • Familiarity with GPU performance analysis: memory hierarchy, utilization, roofline modeling, and profiling with Nsight Systems or Nsight Compute. • Strong written and verbal communication skills, with the ability to present technical findings clearly to both engineering teams and leadership - and to navigate ambiguous, open-ended customer problems. Ways to Stand Out from the Crowd: • Experience with NVIDIA Dynamo or other disaggregated inference serving frameworks. • Contributions to open-source inference or ML systems projects, particularly vLLM or SGLang - please include links to relevant pull requests or artifacts. • Background with ML compilers or GPU kernel development (Triton, CUTLASS, TorchInductor). • Experience building developer tools or internal platforms that meaningfully improved team productivity. • Prior experience in a customer-facing or forward-deployed engineering capacity within a technical product organization. Widely considered to be one of the technology world's most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/ #LI-Hybrid Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 135,000 CAD - 185,000 CAD for Level 3, and 170,000 CAD - 220,000 CAD for Level 4. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until April 14, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Information Technology Jobs

Find similar Senior Software Engineer, AI Inference jobs: