NVIDIA Corporation

Distinguished Engineer, Scaled Out Inferencing

NVIDIA Corporation$320K — $488K *
Enterprise Technology
15+ years of experience
Job Overview by Ladders

Qualifications

  • 16+ years in technical roles focused on AI infrastructure and large-scale inference orchestration.
  • 7-10+ years of leadership experience.
  • BS/MS or higher in systems/software engineering or related fields.
  • Proficient in GPU architecture and low-level performance tuning with CUDA and cloud-native architectures.
  • Demonstrated success delivering complex solutions with operational transparency.
  • Strong technical leadership skills, capable of aligning diverse teams and stakeholders.
  • Excellent collaboration and communication skills to engage with partners and customers.

Responsibilities

  • Architect high-throughput, low-latency distributed inference systems for AI workloads.
  • Drive hardware-software co-optimization and performance tuning at the kernel level.
  • Influence open-source projects to enhance AI inferencing on NVIDIA hardware.
  • Lead the strategy for automated deployment, versioning, and intelligent scaling of models.
  • Engage with stakeholders to ensure NVIDIA's solutions set industry standards.
  • Oversee full software and system lifecycle from ideation to operational management.

Benefits

  • Opportunity to work with cutting-edge AI technology at NVIDIA.
  • Exposure to high-level technical decision-making and strategic planning.
  • Collaboration with leading figures in the AI and computing industries.
  • Engagement with open-source projects and ecosystems.
  • High-impact role influencing the future of AI inferencing.
Full Job Description
As a technology leader at NVIDIA, you will lead the development of our global strategy for scaled-out AI inferencing. You will architect the high-throughput, low-latency distributed pipelines and model serving strategies required for massive scale and production reliability. You will define and drive the technical roadmap for full-lifecycle, from deployment and versioning to automated scaling, across enterprise and cloud environments. Working with NVIDIA leadership, you will establish the systems and orchestration layers that enable the world's most advanced AI models to run with peak efficiency on our accelerated computing hardware.

What You'll Be Doing:
  • Various Architectural Work: Architect distributed pipelines, define and drive the technical implementation of high-throughput, low-latency, distributed inference systems to support massive-scale AI workloads.
  • Collaborate on Cross Domain Disciplines: Hardware-software co-optimization, drive performance tuning at the kernel and driver level, optimizing GPU resource management and hardware acceleration for production-grade model serving.
  • Collaborate on Open Source and Ecosystem Projects: Guide and influence open source projects Dynamo, TensorRT-LLM, and ecosystem projects (vLLM, SGLang, Linux, Kubernetes, Ray) to bring state of the art inferencing on NVIDIA accelerated hardware.
  • Accelerate Integration: Orchestrate model lifecycles, lead the strategy for full-lifecycle model management, including automated deployment, versioning, and intelligent scaling across varied cloud and datacenter environments.
  • Engage Stakeholders: Collaborate with customers, infrastructure providers, and partners to ensure NVIDIA's solutions set the industry standard for performance and availability.
  • Full Software and System Lifecycle: From ideation to architecture, design, development, deployment, operations, and full lifecycle management, lead all technical aspects of planning and continuous evolution of a large technical scope.


What We Need to See:
  • 16+ overall years in technical roles with a recent long-term focus on AI infrastructure and more recent direct experience in large-scale inference orchestration. Proven track record building secure, highly available, and durable production distributed systems.
  • 7-10+ years of leadership experience
  • BS/MS or higher or equivalent experience in systems / software engineering, or related engineering fields
  • Deep Technical Expertise: Proficiency in GPU architecture, hardware acceleration, and low-level performance tuning (CUDA, kernels) alongside cloud-native architectures for multi-tenant model serving.
  • Proven success delivering high-impact technically complex solutions that achieve high levels of transparency into resource utilization, performance, and operational insights.
  • Technical Leadership: Develop and advance consensus and organizational alignment across technical leadership and the highest level of senior corporate leadership. Ability to synthesize cross-functional needs into architecture and design while guiding internal execution across diverse teams.
  • Communication and Teamwork: Strong collaboration and influence skills, capable of leading engineering engagement, communicating with peers, partners, and working with high performance and accelerated computing customers.


Ways to Stand Out from the Crowd:
  • Application of Artificial Intelligence: Real world experience building the systems to support AI/ML workloads.
  • Industry Expertise: Direct experience in designing, developing, delivering and operating secure, highly available, scaled out systems in enterprise and cloud environments.
  • Engineering Enablement: Demonstrated history of creating scalable processes and extensible systems that facilitate cross-functional collaboration and operations at scale.
  • Open Source Collaboration: Familiarity with open source ecosystems and projects (e.g. Dynamo, TensorRT-LLM, vLLM, SGLang, Ray). Ability to collaborate and influence in open source project governance to represent NVIDIA, customers, and partners interests in technical alignment and direction.


Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 320,000 USD - 488,750 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 15, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Enterprise Technology Jobs

Find similar Distinguished Engineer, Scaled Out Inferencing jobs: