DPU RAS and Debug Architect, Infrastructure Silicon

Meta

• $175K — $210K *
Telecommunications & Hardware
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent experience
  • 8+ years in architecting RAS and debug architectures for NIC/DPU or similar ASICs
  • Proficient in RAS concepts including FIT-rate estimation and error detection
  • Experience with error-protection mechanisms like ECC and data poisoning
  • Familiarity with memory and interface RAS, particularly LPDDR/DDR and PCIe
  • Knowledge of on-chip debug and trace architectures, including JTAG
  • Experience with performance-monitoring architectures and DFT concepts

Responsibilities

  • Own the RAS and debug architecture for DPUs, overseeing error detection and reporting
  • Set FIT-rate targets based on usage and deployment models
  • Define error detection, correction, and containment architecture
  • Establish trace, debug, and performance-monitoring architecture for post-silicon debug
  • Decide the division of responsibilities between hardware, firmware, and software
  • Collaborate with design, DV, and PD teams on feature definition and architecture refinement
  • Support integration and interface resolution during post-silicon bring-up

Benefits

  • Comprehensive health insurance plans
  • Retirement savings options with company matching
  • Generous paid time off and holiday schedule
  • Opportunities for professional development and training
  • Flexible work arrangements and remote work options
Full Job Description
Meta's Infrastructure Silicon organization designs custom silicon that powers our data center infrastructure - SmartNICs/IPUs/DPUs, AI accelerators, and networking ASICs. We are looking for a Data Processing Unit (DPU) RAS and Debug Architect to define the reliability, availability, and serviceability architecture for our DPU ASICs. You'll set FIT-rate targets from the intended usages and deployment models, define how errors are detected, corrected, and reported, and define the trace, debug, and performance-monitoring architecture that supports both post-silicon debug and software and firmware debugging. You'll decide what belongs in hardware versus firmware versus software, define how the hardware presents itself to firmware and software, and work hand in hand with the firmware, driver, and RTL teams - carrying your designs from early path-finding through implementation and silicon bring-up.

Responsibilities

Own the RAS and debug architecture for our DPUs - error detection, correction, and reporting, FIT budgeting, and the trace, debug, and telemetry infrastructure - from early path-finding through silicon bring-up
• Set FIT-rate targets from the intended usages and deployment models, and budget them across the design - memories, logic, on-chip interfaces, and links
• Define the error detection, correction, reporting, and containment architecture -- parity, ECC/SECDED, poisoning and poison propagation, error logging, and how errors are surfaced to firmware, host software, and platform management
• Define the trace, debug, and performance-monitoring architecture for post-silicon debug and for software and firmware debugging - on-chip trace, event and counter telemetry, crash and state capture, and the JTAG/debug-access model
• Decide what belongs in hardware versus firmware versus software, and define how the hardware presents itself to them - register and programming models, error and interrupt models, and trace/telemetry interfaces
• Work closely with design, DV, and PD teams on feature definition and PPA tradeoff; refine architecture to meet the design and PD constraints. Define architecture to maximize DV complexity including defining specific features to help ease DV
• Support Design, DV, PD and DFT teams to resolve interface and integration issues as they come up. Support post-silicon bring-up

Minimum Qualifications
• Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
• 8+ years of relevant industry experience architecting RAS and debug/trace architectures and their hardware/software interfaces for NIC/DPU or comparable ASICs
• Experience with RAS concepts: FIT-rate estimation and budgeting, failure modes (including silent data corruption), error detection, correction, and containment, and reliability targets for data center deployments. Familiarity with data center reliability, serviceability, and manageability requirements
• Experience with error-protection mechanisms - parity, ECC/SECDED, CRC, data poisoning and poison propagation, and lockstep/redundancy techniques - and their area, latency, and power trade-offs
• Experience with memory and interface RAS -- LPDDR/DDR RAS (ECC, on-die ECC, error scrubbing, post-package repair) and PCIe RAS (Advanced Error Reporting, ECRC, link error detection and recovery)
• Experience with error reporting and handling architectures (ARM and/or x86) - error logging and registers, interrupts, machine-check and AER-style reporting, and escalation to firmware, host software, and platform/BMC management
• Experience with on-chip debug and trace architectures -- JTAG and debug access, on-chip trace, breakpoints and watchpoints, and crash and state capture (e.g., ARM CoreSight or comparable)
• Experience with performance-monitoring architectures - hardware performance counters and event telemetry - and their use in post-silicon and software/firmware debug
• Experience with DFT concepts (scan, MBIST/LBIST, boundary scan) and how they interact with RAS and debug
• Experience with processor ISA debug mechanisms (ARM and/or x86) and instruction/execution trace mechanisms such as ETM
• Experience with relevant industry standards and specifications, such as OCP (server, RAS, and telemetry/manageability specifications), JEDEC (LPDDR/DDR), and PCIe
• Experience driving analysis independently and influencing architectural direction through data

Preferred Qualifications
• PhD in Computer Science, Computer Engineering or Electrical Engineering
• 15+ years of relevant industry experience architecting RAS and debug/trace architectures and their hardware/software interfaces for NIC/DPU or comparable ASICs
• Experience with hardware description languages (e.g., SystemVerilog, VHDL) and simulation environments used in ASIC development flows
• Experience designing on-chip trace and debug subsystems (e.g., ARM CoreSight) and the associated post-silicon debug tooling
• Experience defining FIT budgets and reliability targets for hyperscale data center silicon, including soft-error rate (SER) analysis and mitigation
• Familiarity with functional-safety standards (e.g., ISO 26262) and reliability qualification methods
• Experience developing Python-based (or other scripting) automation pipelines for debug, telemetry collection, and data analysis
• Experience designing error-reporting and machine-check/AER architectures and the firmware error-handling flows built on them

Similar Jobs

More Jobs at Meta

More Telecommunications & Hardware Jobs

Find similar DPU RAS and Debug Architect, Infrastructure Silicon jobs: