NVIDIA Corporation

Senior HPC AI Cluster Engineer

NVIDIA Corporation$176K — $333K *
US-AnywhereRemote in Santa Clara, CA
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Degree in Computer Science, Engineering, or related field with 8+ years of experience
  • Expertise in HPC and AI technologies encompassing CPUs, GPUs, and fast interconnects
  • Proficient with job scheduling and orchestration tools, especially Slurm, K8s
  • Advanced knowledge of Windows and Linux networking and security protocols
  • Experience with diverse storage solutions like Lustre and GPFS
  • Strong programming skills in Python and bash scripting
  • Familiarity with automation tools like Jenkins, Ansible, Puppet/Chef
  • Deep understanding of networking protocols such as InfiniBand and Ethernet

Responsibilities

  • Design, implement, and maintain large-scale HPC/AI clusters with comprehensive monitoring solutions
  • Manage Linux job schedules and utilize orchestration tools for workload efficiency
  • Develop and sustain continuous integration and delivery pipelines for software deployment
  • Automate deployment and management of large-scale infrastructure environments
  • Deploy monitoring solutions for servers, networks, and storage systems
  • Troubleshoot systems from the bare metal level to application software
  • Document and define standard methodologies for team-wide knowledge sharing
  • Support R&D initiatives and engage in POCs/POVs for innovative improvements

Benefits

  • Equity opportunities as part of compensation
  • Access to cutting-edge technologies and projects
  • Collaborative environment with scientific researchers and developers
  • Focus on continuous learning and professional development
  • Potential for long-term career growth in a leading tech company
Full Job Description
NVIDIA is looking for an experienced HPC-AI Engineer to join the Networking Clusters Solutions Infrastructure team. we are focused on building supercomputers and AI clusters based on groundbreaking technologies. We are looking for an outstanding engineer, be a key player to the most exciting computing hardware and software to contribute to the latest breakthroughs in artificial intelligence and GPU computing. Provide insights on at-scale system design and tuning mechanisms for large-scale compute runs. You will work with the latest Accelerated computing and Deep Learning software and hardware platforms, and with many scientific researchers, developers, and customers to craft improved workflows and develop new, leading differentiated solutions. You will interact with HPC, OS, GPU compute, and systems specialist to architect, develop and bring up large scale performance platforms.

What you will be doing:
  • Design, implement and maintain large scale HPC/AI clusters with monitoring, logging and alerting
  • Manage Linux job/workload schedules and orchestration tools
  • Develop and maintain continuous integration and delivery pipelines
  • Develop tooling to automate deployment and management of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resources
  • Deploy monitoring solutions for the servers, network and storage
  • Perform troubleshooting bottom up from bare metal, operating system, software stack and application level
  • Being a technical resource, develop, re-define and document standard methodologies to share with internal teams
  • Support Research & Development activities and engage in POCs/POVs for future improvements


What we need to see:
  • A degree in Computer Science, Engineering, or a related field (or equivalent experience) and 8+ years of experience
  • Knowledge of HPC and AI solution technologies from CPU's and GPU's to high speed interconnects and supporting software
  • Experience with job scheduling workloads and orchestration tools such as Slurm, K8s
  • Excellent knowledge of Windows and Linux (Redhat/CentOS and Ubuntu) networking (sockets, firewalld, iptables, wireshark, etc.) and internals, ACLs and OS level security protection and common protocols e.g. TCP, DHCP, DNS, etc.
  • Experience with multiple storage solutions such as Lustre, GPFS, Weka.io. Familiarity with newer and emerging storage technologies.
  • Python programming and bash scripting experience.
  • Comfortable with automation and configuration management tools such as Jenkins, Ansible, Puppet/chef
  • Deep knowledge of Networking Protocols like InfiniBand, Ethernet
  • Deep understanding and experience with virtual systems (for example VMware, Hyper-V, KVM, or Citrix)
  • Familiarity with cloud computing platforms (e.g. AWS, Azure, Google Cloud)


Ways to stand out from the crowd:
  • Knowledge of CPU and/or GPU architecture
  • Knowledge of Kubernetes, container related microservice technologies
  • Experience with GPU-focused hardware/software (DGX, Cuda)
  • Experience with RDMA (InfiniBand or RoCE) fabrics


Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 176,000 USD - 276,000 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 24, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Information Technology Jobs

Find similar Senior HPC AI Cluster Engineer jobs: