NVIDIA Corporation

Senior Site Reliability Engineer - HPC

NVIDIA Corporation$152K — $287K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • B.S. in Computer Science or related field, or equivalent experience; 5+ years in critical services support.
  • Experience with large-scale HPC clusters, including Slurm, LSF, or Kubernetes.
  • Proficient in CI/CD and Infrastructure as Code (IaC) practices.
  • Strong track record in creating scalable infrastructure for automated host management and observability.
  • 5+ years of coding in languages like Python and Go, with a focus on scripting.
  • Previous mentorship experience and ability to influence technical direction.

Responsibilities

  • Lead end-to-end SRE solutions from design to operational improvement.
  • Automate provisioning using Infrastructure as Code and configuration management.
  • Deliver services across multi-cloud environments including AWS and GCP.
  • Design infrastructure with redundancy for high availability and fault tolerance.
  • Ensure maximum uptime and quality service for internal stakeholders.
  • Plan for capacity to meet ongoing operational demands.
  • Identify performance issues and propose effective solutions.
  • Collaborate across teams to ensure successful project execution.
  • Participate in incident reviews and produce quality root cause analysis reports.

Benefits

  • Comprehensive benefits package with options for health, wellness, and flexible work arrangements.
  • Opportunity for equity participation.
  • Access to cutting-edge technology and peer collaboration with industry-leading experts.
  • A diverse and inclusive work environment, supporting personal and professional growth.
Full Job Description
We’re looking for a Senior SRE to join our Compute Farm team and help build the next generation of our global services platform. At NVIDIA, you’ll keep critically important systems running while working on the technologies that are redefining computing. You’ll harness the power of AI to deliver groundbreaking solutions to some of the world’s toughest problems—and see your work have real, lasting impact! What you'll be doing: • Own SRE solutions end‑to‑end, from design and implementation to operation and continuous improvement, ensuring they integrate cleanly with HPC schedulers, storage, and network fabrics. • Use IaC(Infrastructure‑as‑Code) and config management to standardize and automate provisioning everywhere. • Deliver solutions in a globally distributed, multi‑cloud hybrid environment – On‑prem, AWS, GCP, and OCI. • Design for failure with redundancy, failure domains, progressive delivery, and strict change control. • Ensure the highest level of uptime and Quality of Service (QoS) for internal customers through operational excellence. • Conduct capacity management and planning to meet ongoing operational needs. • Detects performance issues and recommends solutions to maintain world‑class service quality. • Collaborate with various teams in a fast‑paced environment to ensure seamless project completion. • Participate in on-call, incident reviews, assist in root cause identification, and produce high-quality RCA reports. What we need to see: • B.S. degree in Computer Science or related technical field (or equivalent experience) with 5+ years professional experience building and supporting critical services. • Experience supporting large‑scale HPC clusters using Slurm, LSF or Kubernetes clusters, including setup, tuning, and troubleshooting. • Proficiency in modern CI/CD techniques, and Infrastructure as Code (IaC) for managing services. • Strong experience crafting large-scale infrastructure platforms for automated host lifecycle management, fleet reliability/auto-healing, E2E observability or data-driven operations (AIOps/ML-driven signals) that materially reduce manual intervention. • Proficient in monitoring, metrics, container management, and log collection tools. • 5+ years of coding/scripting experience in at least two high‑level programming languages such as Python, Go, Perl, or Ruby. • Mentored other engineers and influenced technical direction through design reviews, architecture documents, and strong partnership with product and leadership. • Creative problem solver with excellent debugging skills and strong communication and documentation abilities. Ways to stand out from the crowd: • Published technical write‑ups or talks (conference presentations, meetups, engineering blogs) that deep‑dive into real‑world reliability, observability, or large‑scale HPC/SRE problems and their solutions. • Maintainer or co‑maintainer responsibilities for an open source component used in production (plugins, operators, exporters, controllers, or SDKs) at large scale. Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. We have some of the most brilliant and talented people in the world working for us and, due to unprecedented growth, our world-class engineering teams are growing fast. If you're a creative and autonomous engineer with real passion for technology, we want to hear from you. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4. You will also be eligible for equity and . Applications for this job will be accepted at least until September 8, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Information Technology Jobs

Find similar Senior Site Reliability Engineer - HPC jobs: