NVIDIA Corporation

Senior SoC Architect, RAS

NVIDIA Corporation$184K — $356K *
Telecommunications & Hardware
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • MS or PhD in computer engineering, electrical engineering, or equivalent experience.
  • 8+ years of SoC architecture, design, verification, reliability, or silicon validation experience.
  • Strong understanding of Reliability, Availability, and Serviceability (RAS) in the SoC context.
  • Proven track record in defining and driving hardware architecture features through full development lifecycle.
  • Solid grasp of SoC architecture and its interaction with overall system performance.
  • Industry expertise in areas such as RAS, safety, debug, memory controllers, and IO technologies.
  • Hands-on experience with design verification and reliability validation methodology.
  • Familiarity with radiation effects and high-reliability environments is valued.

Responsibilities

  • Define and drive SoC-level RAS hardware architecture across various components.
  • Own RAS features from concept to production readiness, covering all phases.
  • Develop architectural requirements for fault management and system diagnostic features.
  • Collaborate with multiple teams to ensure implementable and verifiable RAS features.
  • Analyze how RAS mechanisms interact with overall SoC performance and design flows.
  • Create hardware specifications, design guidance, test plans, and models in relevant languages.
  • Plan and review verification and validation strategies for RAS mechanisms, including error handling.
  • Assist in silicon debug and failure analysis across various testing environments.

Benefits

  • Opportunity to impact cutting-edge technology applications.
  • Collaboration with world-class experts across various domains.
  • Dynamic and technology-focused work environment.
  • Involvement in advanced projects related to AI and high-reliability systems.
  • Potential for patenting innovative hardware architecture techniques.
Full Job Description
We are now looking for a Senior Hardware Architect for our Tegra System-on-Chips (SoC) focused on Reliability, Availability, and Serviceability (RAS). Do you want to be part of the Artificial Intelligence (AI) revolution and help define resilient computing platforms for datacenters, autonomous vehicles, edge systems, and other high-reliability applications? We are looking for an exceptional SoC architect to help define, drive, and deliver RAS hardware architecture across advanced CPUs and SoCs, from early architectural concepts through design implementation, verification, validation, and production readiness.

This position offers the opportunity to have real impact in a dynamic, technology-focused company developing state-of-the-art processor and system architectures at the forefront of machine learning, autonomous vehicles, high-performance computing, and edge computing. You will work with world-class systems architects, RAS experts, design teams, verification teams, validation teams, firmware teams, and software partners to define end-to-end hardware RAS features that improve system resiliency, observability, debuggability, error containment, recovery, and serviceability. Space and radiation-aware design are important areas of interest for this role, including understanding how radiation effects can influence SoC reliability, but the primary focus is broad SoC RAS architecture and driving features successfully through the product development flow.

What you'll be doing:
  • Define and drive SoC-level RAS hardware architecture across CPUs, interconnects, memory systems, IOs, safety islands, firmware interfaces, and platform-level components.
  • Own RAS features from concept through architecture specification, micro-architecture alignment, RTL implementation support, design verification, silicon validation, debug, and production readiness.
  • Develop architectural requirements for fault detection, correction, containment, isolation, telemetry, error reporting, recovery, graceful degradation, serviceability, and diagnostic observability.
  • Work closely with design, verification, validation, firmware, software, and platform teams to ensure RAS features are implementable, verifiable, debuggable, and aligned with system-level requirements.
  • Understand the broader SoC architecture and identify how RAS mechanisms interact with performance, power, reset flows, clocks, memory hierarchy, interconnect behavior, firmware-visible controls, and platform software.
  • Create hardware specifications, architectural requirements, error-handling flows, design guidance, test plans, and architectural models in SystemC, C/C++, Python, or other relevant modeling environments where applicable.
  • Plan and review verification and validation strategies for RAS mechanisms, including error injection, recovery validation, coverage analysis, resiliency modeling, and cross-functional architecture reviews.
  • Assist in failure analysis and silicon debug for lab, post-silicon, production, and field findings; develop diagnostic screens and localization methods for latent, intermittent, and environment-sensitive failures.
  • Apply RAS architecture principles to high-reliability deployment environments, including space-aware and radiation-aware use cases where single-event effects, memory corruption, logic corruption, or cumulative radiation exposure may impact system reliability.
  • Follow industry standards and best practices related to RAS, functional safety, semiconductor reliability, debuggability, verification, validation, and silicon testing.
  • Patent novel hardware architecture techniques that improve system resiliency, observability, serviceability, and recovery.


What we need to see:
  • MS or PhD degree in computer engineering, electrical engineering, or equivalent experience.
  • At least 8+ years of SoC architecture, design, verification, reliability, silicon validation, or related hardware development experience.
  • Strong understanding of Reliability, Availability, and Serviceability (RAS) in the SoC context, including fault detection, correction, containment, telemetry, recovery, degradation modes, debug visibility, and serviceability mechanisms.
  • Experience defining and driving hardware architecture features through the full development lifecycle, including architecture definition, design implementation, verification planning, validation, debug, and production readiness.
  • Strong understanding of overall SoC architecture and the ability to reason across micro-architecture, full-chip integration, firmware interfaces, software-visible behavior, platform flows, and customer use cases.
  • Meaningful industry expertise in one or more SoC architecture areas such as RAS, safety, debug, clocks, resets, interconnects, memory controllers, IO technologies, platform integration, firmware-visible error handling, or diagnostic infrastructure.
  • Hands-on experience with design verification, silicon validation, fault injection, coverage analysis, resiliency modeling, diagnostic development, or reliability validation methodology.
  • Familiarity with radiation effects, space operation, or other high-reliability deployment environments is strongly valued, including understanding how hardware architecture can mitigate single-event effects and related reliability risks.
  • Excellent analytical, written, and verbal interpersonal skills with the ability to work effectively across architecture, design, verification, firmware, software, validation, and customer-facing teams.


Ways to stand out from the crowd:
  • Demonstrated history of architecting and delivering complex SoC RAS features across design, verification, validation, and production phases.
  • Deep familiarity with architectural resiliency techniques such as ECC, parity, redundancy, replay, checkpoint/restart, scrubbing, isolation, containment, telemetry, error logging, recovery flows, and graceful degradation.
  • Experience with cross-functional debug of hardware failures in simulation, emulation, post-silicon validation, production, customer deployments, or other high-reliability systems.
  • Familiarity with Design for Debug, Design for Test, Design for Reliability, silicon observability, fault-injection methodology, and coverage-driven validation flows.
  • Experience with radiation effects analysis, soft-error-rate analysis, radiation test campaigns, accelerated stress testing, heavy-ion or proton testing, or space qualification methodology.


Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 6, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

About NVIDIA Corporation

Nvidia, a global leader in graphics, gaming, and AI technology, offers Nvidia careers and internship opportunities for those passionate about driving innovation in the tech industry. you'll find a company committed to growth, teamwork, and leadership in computer science and machine learning domains.

About Nvidia

A Pioneer in Technology and Innovation

Nvidia has cemented its reputation as a powerhouse in developing advanced graphics processing units (GPUs) and has significantly contributed to the gaming industry's evolution. Moreover, its foray into AI and machine learning has opened new frontiers in technology, making Nvidia a beacon of innovation and a desirable workplace for ambitious tech professionals.

Job Opportunities

Diverse Positions in a Dynamic Field

Nvidia is continuously on the lookout for talented individuals across various domains, including hardware and software engineering, product design, marketing, and sales. Employment opportunities at Nvidia are vast, catering to a wide range of expertise and career aspirations.

Employment in Hardware and Graphics

For those fascinated by the intricacies of hardware and graphics technology, Nvidia offers positions that sit at the forefront of gaming and computing advancements.

Growth in Machine Learning and AI

Nvidia's leadership in AI and machine learning has created numerous vacancies for specialists eager to contribute to groundbreaking projects.

Recruitment in Computer Science

With the constant demand for innovation, Nvidia's recruitment efforts focus on computer science experts capable of pushing the boundaries of what's possible.

Internship Program

Opening Doors to Future Innovators

Nvidia's internship program is designed to nurture the next generation of technology leaders, offering hands-on experience in a culture that celebrates creativity and teamwork.

Benefits and Culture

Interns at Nvidia enjoy a plethora of benefits, from competitive stipends to mentorship opportunities, all within an environment that values growth and learning.

Opportunities for Students

Whether you're an undergraduate, a master's student, or a Ph.D. candidate, Nvidia's internships provide a real-world glimpse into the tech industry, offering valuable experience in various technology fields.

Pathways to Full-Time Employment

Many interns have transitioned into full-time positions, marking the start of successful careers at Nvidia. The internship program is more than a stepping stone into the company; it’s an investment in the professional development of interns. The goal is to ensure that interns are well-equipped for future challenges.

Nvidia Careers: More Than Just a Job

Nvidia offers more than just a job to its employees; it provides a front-row seat on the journey into the future of technology. Nvidia stands as a pillar of innovation with its vast opportunities in hardware, graphics, gaming, machine learning, and computer science. Nvidia careers serve as a launching pad for talented workers who aim to redefine the technological landscape. Whether through full-time positions or internships, joining Nvidia means contributing to a legacy of breakthroughs and becoming part of a global community dedicated to pushing the boundaries of what's possible.
Learn more about NVIDIA Corporation
Size
22,473 employees
Market Cap
$350.4 billion
Industry
Net Income
$4.3 billion
Founded
1993
5 Year Trend
+31.3%
Revenue
$16.6 billion
NASDAQ

Similar Jobs

More Jobs at NVIDIA Corporation

More Telecommunications & Hardware Jobs

Find similar Senior SoC Architect, RAS jobs: