NVIDIA Corporation

Senior Storage Software Engineer - DGX Cloud

NVIDIA Corporation$224K — $431K *
Enterprise Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or related field - or equivalent experience
  • 12+ years of direct experience in storage software engineering with high-performance parallel or distributed file systems
  • Contributions to open-source projects involving distributed or parallel file systems
  • Experience diagnosing and resolving storage issues in large-scale GPU or HPC clusters
  • Strong proficiency in systems languages (C, C++, Rust, or Go) and Python; comfortable with Linux kernel storage and networking
  • Strong communication skills for conveying complex technical trade-offs
  • Comfort operating in a 24/7 production environment with a security-first approach

Responsibilities

  • Contribute code to open-source parallel and distributed file systems and engage with upstream communities
  • Write and review production code; analyze kernel and storage system source code for bug resolution
  • Triage and troubleshoot complex storage issues across large GPU clusters
  • Validate storage architecture and performance through scale tests and benchmarks
  • Recommend configuration and tuning best practices for high-performance file systems
  • Collaborate with various teams and external partners on storage architecture
  • Utilize modern AI tools for building, debugging, validation, and operations

Benefits

  • Equity and benefits package
  • Opportunity to work on foundational storage engineering for AI
  • Engagement with high-impact projects across cloud and on-prem setups
  • Access to cutting-edge tools and technologies in accelerated computing
  • Collaboration with experts in AI and GPU computing
Full Job Description
NVIDIA DGXC Storage team handles some of the fastest training and inference tasks. Every GPU cycle depends on a storage platform built to keep tens of thousands of accelerators continuously busy. It maintains exabytes of data securely and powers the largest AI workloads worldwide across cloud, neocloud, and on-prem setups. With the growth of accelerated computing, storage is essential. It can make the difference between effective GPU use and wasted potential, and between launching a frontier model on time or missing the deadline by months. Were looking for a hands-on Storage Software Engineer to join the storage team as an individual contributor and technical lead. You will contribute to open-source parallel and distributed file systems and keep our largest GPU clusters fast, reliable, and durable. You will stay deeply hands-on: writing and reviewing production code, chasing root causes in the field, and setting the configuration and tuning standards our GPU fleets run on. This is a chance to do foundational storage engineering for the AI era at the company that introduced accelerated computing. What youll be doing: 3 Contribute to open-source file systems. Contribute code to open-source parallel and distributed file systems, and distributed object storage. Upstream fixes and features, and engage directly with the upstream communities and maintainers. 3 Serve as a hands-on storage software lead. Write and review production code yourself, and read kernel, NFS, NVMe-oF, or SPDK source when a bug requires it. Make the final technical calls on storage deliveries against measurable targets. 3 Triage and troubleshoot at scale. Triage, troubleshoot, and root-cause large, complex storage issues across very large GPU clusters (tens of thousands of GPUs) - I/O and metadata performance, data corruption, and recovery. 3 Validate architecture and capabilities. Validate storage architecture, capabilities, performance, and durability. Run scale tests, benchmarks, and recovery drills, and qualify new builds against measurable performance and durability targets. 3 Recommend configuration, tuning, and guidelines. Define and recommend configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure, and help operators and internal customers apply them. 3 Partner broadly. Work with training, inference, and accelerated-computing teams, site-reliability and operations, networking, and security, and collaborate with cloud providers, neocloud operators, and storage vendors on a common architecture. 3 Work AI-first. Use modern AI coding and agentic tools day-to-day to accelerate building, debugging, validation, and operations. What we need to see: 3 BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field - or equivalent experience. Over 12 years of direct experience in storage software engineering, including extensive involvement with a high-performance parallel or distributed file system handling multi-petabyte scale. 3 Contributions to open-source projects involving a distributed or parallel file system. You are fully engaged in engineering tasks. You write and review production code, examine file system, kernel, NVMe-oF, or SPDK source to identify bugs, and personally conduct scale tests or recovery drills instead of assigning them to others. 3 Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance. 3 Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python; comfortable in the Linux kernel storage and networking stacks (block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, multipath). 3 Working knowledge of object storage (S3 / Swift-class) and block storage (NVMe-oF, iSCSI). 3 Strong written and verbal communication; capable of clarifying complex technical trade-offs to engineers, SREs, vendors, and internal customers. 3 Comfort operating in a 24/7 production environment where storage incidents directly impact GPU availability, with a security-first approach baked into every build. 3 100% hands-on engineering. You write and review production code, read file system, kernel, NVMe-oF, or SPDK source to chase bugs, and run scale tests or recovery drills yourself rather than delegating. Ways to stand out from the crowd: 3 Maintainers or sustained contributions to widely used public projects. 3 Experience crafting or operating storage for AI training or inference at very large GPU scale, with measurable gains in GPU utilization or reductions in I/O bottlenecks. 3 Kernel and file system development experience, metadata scalability, data placement, failure recovery, or HSM or equivalent experience. 3 Kubernetes and CSI driver development for storage. 3 Hands-on experience with SPDK, libfabric, or FUSE performance optimization. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until August 2, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Enterprise Technology Jobs

Find similar Senior Storage Software Engineer - DGX Cloud jobs: