The DGX Cloud organization bridges customer success and cloud infrastructure engineering, partnering directly with NVIDIA's internal research and product teams to accelerate AI workload development. As a Customer Success Engineer, you'll embed deeply with internal customers - gaining a thorough understanding of their applications and translating that knowledge into architectural guidance, best practices, and hands-on solutions. This role sits at the unique intersection of solutions architecture and platform strategy: you'll write code, build tooling, and help shape NVIDIA's GPU capacity management from the inside. Working across Engineering, Product, Finance, and Operations, you'll connect infrastructure roadmaps to business needs in a way that directly influences how NVIDIA's most advanced AI teams move faster. If you thrive where deep technical work meets high-stakes collaboration, this role was built for you.
What you'll be doing:- Design and implement distributed cloud infrastructure at scale - spanning compute, storage, networking, and GPU capacity management across IaaS, PaaS, and SaaS models zendesk
- Partner with internal research and product teams to understand workloads from both a technology and business perspective, providing architectural guidance that drives their success
- Contribute code directly when needed to move projects forward, and codify working patterns into tools, playbooks, and building blocks that others can reuse
- Build and maintain agentic tooling to automate operational workflows and infrastructure resource management
- Analyze the DGX Cloud ecosystem to understand current customer demand and future capacity needs, driving infrastructure efficiency initiatives in partnership with Engineering, Finance, and Product
- Present technical roadmaps, architecture decisions, and demos to internal stakeholders and NVIDIA leadership, driving cross-functional consensus on infrastructure strategy
What we need to see:- BS or MS in Computer Science, Engineering, or a related field, or equivalent experience.
- 12+ years of experience designing and building distributed systems and cloud infrastructure, with demonstrated experience in GPU capacity management for high-performance computing
- Demonstrated ability to write production code in Golang, Java, C, C++, Python, or Rust
- Experience with Kubernetes and/or distributed task scheduling
- Strong background in Infrastructure, Networking, Storage, and DevOps scripting/tooling
- Experience deploying AI/ML workloads at scale
- Strong communication and relationship-building skills, with a demonstrated ability to drive cross-functional consensus and align stakeholders across departments
#LI-Remote
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 200,000 USD - 322,000 USD.
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until August 14, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.