NVIDIA Corporation

Distinguished Engineer, Production Engineering, Data Center Automation

NVIDIA Corporation$320K — $488K *
Enterprise Technology
15+ years of experience
Job Overview by Ladders

Qualifications

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or related field, or equivalent experience
  • 18+ years of experience in large-scale distributed systems or production environments
  • Company-level technical leadership experience in production engineering or cloud platforms
  • Experience in establishing operating models and engineering standards across technical domains
  • Proven track record in leading cross-team technical efforts from concept to production.

Responsibilities

  • Define the long-range technical strategy for operating DGX Cloud clusters
  • Establish architectural vision and operational guidelines for cluster lifecycle management
  • Guide the roadmap for critical investments improving production readiness and operational safety
  • Make high-impact technical decisions for coordinating platform and service teams
  • Develop workflows and engineering collaboration across various operational domains.

Benefits

  • Equity opportunities
  • Comprehensive health insurance packages
  • Flexible work arrangements
  • Professional development support
  • Collaborative and innovative work environment.
Full Job Description
NVIDIA is looking for a Distinguished Engineer to act as a senior technical leader in the Production Engineering group, enthusiastic about cluster operations involving DGX Cloud GPU capacity.

This role centers on the operational framework supporting DGX Cloud environments spanning on-premises, major cloud providers, and NVIDIA Cloud Partner locations. The responsibilities include engineering integrations to guarantee DGX Cloud capacity is fully operational in production. This involves Kubernetes service management, ensuring vendor and equipment availability, on-prem infrastructure operations, release and runtime preparation, and service reliability coordination. The workflows tie these elements into a cohesive production system.

This hands-on Distinguished Engineer role calls for a deeply technical leader to build the architectural direction for cluster operations in DGX Cloud. The ideal candidate will blend software engineering expertise, system knowledge, and production insight to define technical strategy, set operating standards, direct the evolution of the production model, and drive delivery of cross-organizational capabilities. These capabilities ensure that the DGX Cloud resources remain usable, maintainable, and improve continuously at scale. The position demands both deep invention and implementation skills and the ability to lead by influence across several teams and critical production results.

What you'll be doing:
  • Define the long-range technical strategy for operating DGX Cloud clusters consistently across on-prem, hyperscalers, and NeoCloud environments
  • Define the architectural vision and core operational guidelines for cluster lifecycle, runtime delivery, restoration, release readiness, and steady-state operability throughout DGX Cloud resources
  • Guide the roadmap and execution of critical cross-organizational investments that improve production readiness, operational safety, performance, and cross-team coordination
  • Make and guide high-impact technical decisions that resolve how platform, hardware, provider, and service teams coordinate to operate DGX Cloud resources in production
  • Develop robust workflows, interfaces, and engineering collaboration across Kubernetes production service, provider and hardware readiness, on-prem and bare-metal infrastructure operations, and service-layer reliability domains


What we need to see:
  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent experience
  • 18+ years of experience building and operating large-scale distributed systems, infrastructure platforms, or production environments
  • Confirmed company-level technical leadership at principal, distinguished, or equivalent scope in production engineering, SRE, infrastructure software, or cloud platforms
  • Confirmed experience in establishing operating models, architectural direction, and engineering standards across various technical domains and organizations
  • Consistent record leading large, cross-team technical efforts from concept through production, including aligning collaborators, navigating complexity and delivering measurable outcomes


Ways to stand out from the crowd:
  • You have developed the production operating model for a large, diverse infrastructure environment spanning multiple platforms or providers
  • You have set widely recognized operating standards, architectures, APIs, or workflows that increased reliability, operability, or performance at company scale
  • You have built automation and engineering interfaces that connect platform teams, infrastructure teams, and service owners into a consistent production system


Additional Job Description Mentorship

This role is specifically dedicated to the cross-domain production operating model for DGX Cloud capacity. It does not include the shared platform-software role for common automation services. Success relies on ensuring the wider DGX Cloud cluster estate functions optimally in production with proven technical leadership, consistent workflows, defined boundaries, and effective cross-team engineering collaboration.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 320,000 USD - 488,750 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 6, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Enterprise Technology Jobs

Find similar Distinguished Engineer, Production Engineering, Data Center Automation jobs: