NVIDIA Corporation

Software Engineer, CUDA Deep Learning Systems

NVIDIA Corporation$124K — $195K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or related field.
  • 2+ years of industry or academic experience post-degree.
  • Proficient in C++ and Python programming.
  • Strong foundations in deep learning, particularly with transformers.
  • Understanding of distributed computing and performance challenges in large-scale systems.
  • Experience in systems programming and performance optimization.
  • Familiarity with CUDA programming and deep learning accelerators.

Responsibilities

  • Research and prototype system optimizations for advanced deep learning models.
  • Architect and optimize scalable distributed computing systems.
  • Design and implement high-performance CUDA kernels for neural networks.
  • Analyze hardware-software interactions to resolve performance bottlenecks.
  • Collaborate with researchers and engineers to enhance compute utilization and efficiency.
  • Develop tools and runtime systems for deep learning acceleration.
  • Write clean and maintainable code for seamless transitions of prototypes.

Benefits

  • Eligibility for equity in the company.
  • Access to comprehensive health benefits.
  • Opportunities for professional development and training.
  • Flexible working environment and work-life balance.
  • Involvement in pioneering AI technologies and projects.
Full Job Description
We are looking for an experienced and highly motivated software professional to work on pioneering initiatives and projects at the intersection of CUDA and Deep Learning Systems. As the complexity and scale of artificial intelligence continue to grow, the intersection of advanced deep learning architectures, massive-scale distributed computing, and low-level hardware optimization has never been more critical. Our team is dedicated to exploring and prototyping next-generation ideas that bridge the gap between deep learning algorithms and CUDA, pushing the boundaries of what is possible on modern accelerator architectures. Join our dynamic, research-oriented team to help unlock maximum hardware performance for emerging AI workloads. You will be a crucial member of a highly technical group exploring uncharted territories in model optimization, custom kernel development, and cluster-scale AI systems design. If you are passionate about the fundamentals of deep learning and thrive on squeezing every ounce of performance out of advanced computing systems from a single GPU to supercomputer clusters, we want you on our team!

What you will be doing:
  • Explore, research, and prototype novel systems optimizations for advanced deep learning models at the intersection of high-level DL frameworks and low-level CUDA through modeling, simulation, and silicon prototyping.
  • Architect and optimize distributed computing systems that scale seamlessly from a single node to massive, cluster-scale supercomputing environments.
  • Design, implement, and optimize custom high-performance CUDA kernels tailored to emerging neural network architectures and workloads.
  • Analyze complex hardware-software interactions to identify and resolve performance bottlenecks in both training and inference pipelines.
  • Collaborate closely with AI researchers, HW and SW architects, kernel and compiler authors and CUDA driver experts to co-design systems and algorithms that improve accelerator compute utilization, memory bandwidth, cross-node network communication efficiency and programmability.
  • Develop exploratory tools and runtime systems to profile and accelerate new paradigms in deep learning.
  • Write clean, effective, and maintainable code, ensuring exploratory prototypes can smoothly transition into open-source releases, upstream framework integrations, internal tools, or closed-source commercial products.


What we need to see:
  • BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field (or equivalent experience).
  • 2+ years of relevant industry experience or equivalent academic experience after degree achievement.
  • Strong proficiency in C++ and Python programming.
  • Solid background in the fundamentals of Deep Learning with a focus on transformers.
  • Strong understanding of distributed computing principles, multi-node scaling, and the unique performance challenges of cluster-scale execution.
  • Proven experience in systems programming, computer architecture, and low-level systems performance optimization.
  • Familiarity with deep learning accelerator architectures such as the GPU and hands-on experience with CUDA programming, kernel optimization, and workload profiling
  • Experience profiling and optimizing generative AI models, including but not limited to, pioneering large language models.
  • Research background in machine learning systems or adjacent fields and experience profiling and optimizing innovative vision models, generative AI architectures, or diffusion models.
  • A track-record of initiative and willingness to deep-dive on problems across the stack.


Ways to stand out from the crowd:
  • Deep expertise in performance internals and execution graphs of major deep learning training and inference frameworks (e.g., PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, Megatron).
  • Hands-on experience with communication libraries (e.g., NCCL, MPI, UCX) and distributed machine learning techniques (e.g., pipeline, tensor, expert parallelism).
  • Knowledge of numerical methods and low-precision arithmetic (e.g., NVFP4, MXFP4, FP8, INT8) and their impact on deep learning accuracy and performance.
  • Background in deep learning compilers and ML systems, including graph-level and codegen tools (e.g., Triton, XLA, torch.compile) and highly parallel/RL-style simulation environments.
  • Experience designing and implementing agentic AI systems applied to complex systems and infrastructure problems.


Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 124,000 USD - 195,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 9, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Information Technology Jobs

Find similar Software Engineer, CUDA Deep Learning Systems jobs: