NVIDIA Corporation

Senior Solutions Architect, NVIDIA Cloud Partner Operations

NVIDIA Corporation$224K — $356K *
Enterprise Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • BS, MS, or PhD in Computer Science, Engineering, Physics, Mathematics, or related field; or equivalent experience.
  • 12+ years in production infrastructure or cloud engineering; or 5+ years in large-scale GPU/AI infrastructure.
  • Experience in building and operating distributed infrastructure under production load.
  • Deep expertise in Day 2 operations with large-scale GPU or HPC infrastructure.
  • Hands-on knowledge of Kubernetes, Slurm, and related automation tools.
  • Strong Linux skills and experience with Python, Bash, or similar languages.
  • Proven ability to lead complex projects with cross-functional teams.

Responsibilities

  • Solve Day 2 operations challenges at scale with partner engineers.
  • Prepare operating models for new NVIDIA platforms and drive real-time adoption.
  • Enhance reliability, performance, and economics using quantitative measures.
  • Elevate partners' Day 2 maturity by identifying gaps in processes and tools.
  • Transform validated solutions into standardized operating procedures for other partners.
  • Establish a feedback loop to collect and analyze patterns across partners.

Benefits

  • Eligible for equity opportunities.
  • Comprehensive benefits package.
Full Job Description
NVIDIA is looking for a hands-on Solutions Architect to raise the Day 2 operations bar across our NVIDIA Cloud Partner ecosystem. Day 2 starts when a cluster is installed and validated: keeping the service healthy, adapting it as technology and customer demand change, and improving performance, stability, efficiency and economics over time. You will work with engineers running AI clouds at scale on the problems that decide whether customers stay and whether the next generation of NVIDIA technology lands successfully.

Our job is to work hand in hand with NCPs to solve real problems and drive real optimizations, prove the answer, and turn it into something the next partner can use! This is not an outsourced operations role. The partner owns its cloud; success means leaving its team more capable, not more dependent on ours.

What you'll be doing:
  • Solve hard Day 2 operations problems at scale. Work alongside partner engineers to find the cause, prototype an approach, validate it under representative load, and leave behind a practice their team can operate.
  • Make new technology Day 2 ready. Help partners prepare the operating model for new NVIDIA platforms, capacity, services, and use cases before customers depend on them, and help drive adoption in live environments without degrading service.
  • Improve reliability, performance, and economics together. Use measures such as incident frequency, recovery time, utilization, and cost per token to show where the cloud is losing performance or margin - and whether the fix worked.
  • Raise each partner's Day 2 maturity. Identify and help close the gaps that matter across people, process, tooling, telemetry, security, and incident response.
  • Turn one solution into ecosystem capability. Convert validated work into operating procedures, reference architectures, assessments, automation, and agentic workflows that other NCPs can integrate into their standard operating model.
  • Create the feedback loop only NVIDIA can. Spot patterns across partners early and bring clear field evidence to account teams, support, product, and engineering so repeated problems are fixed at the right level.


What we need to see:
  • BS, MS, or PhD in Computer Science, Electrical or Computer Engineering, Physics, Mathematics, or a related field - or equivalent experience.
  • 12+ years in production infrastructure, cloud engineering, solutions architecture, site reliability engineering, HPC, or a similar technical role; alternatively, 5+ years of exceptional specialist-level work in large-scale GPU or AI infrastructure.
  • Experience building, operating, or improving distributed infrastructure under real production load - not only designing or deploying it.
  • Deep expertise in at least one part of the Day 2 stack, backed by hands-on work with large-scale GPU, HPC, or cloud infrastructure. Relevant technologies may include DCGM, BMC/Redfish, and firmware and driver lifecycle; InfiniBand or high-speed Ethernet, NCCL, and UFM; or high-performance storage such as Lustre, IBM Storage Scale, WEKA, VAST Data, or comparable platforms.
  • Working experience across the broader operating platform, including Kubernetes or Slurm, GPU scheduling and multi-tenancy, Prometheus, Grafana or OpenTelemetry, and automation with Terraform, Ansible, Argo CD, or similar tooling.
  • Strong Linux knowledge and enough Python, Bash, or similar experience to automate measurement, diagnosis, validation, or remediation.
  • A detailed evidence-led approach to troubleshooting across system boundaries, paired with the judgment to make difficult technical findings clear.
  • The ability to lead sophisticated work with partner engineers and cross-functional teams without direct authority or taking ownership away from the operator.
  • Strong communication, prioritization, and time-management skills across multiple partner engagements.


Ways to stand out from the crowd:
  • Real world experience operating a GPU cloud, HPC environment, or large-scale AI platform under customer load.
  • Built or matured a 24/7 operations function, including observability, incident and problem management, coverage, and on-call design.
  • Hands on experience with NVIDIA rack-scale platforms such as GB200 or GB300 NVL72 into production, or have hands-on experience with NVIDIA operations technologies such as Spectrum-X, UFM, Base Command Manager, Mission Control, and the GPU or Network Operators.
  • Driven improved fleet health or unit economics through benchmarking, infrastructure as code, GitOps, automated diagnosis, or agent-based remediation.


Even if your background doesn't match every line above, we'd love to hear how your experience applies.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 16, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Enterprise Technology Jobs

Find similar Senior Solutions Architect, NVIDIA Cloud Partner Operations jobs: