THE ROLEAs a Senior Failure Analysis Engineer, you will play a critical role in ensuring the reliability, performance, and successful deployment of AMD's next-generation GPU-accelerated server platforms and AI infrastructure. Working at the intersection of hardware, firmware, validation, manufacturing, and quality engineering, you will investigate complex system- and rack-level failures across high-performance compute environments and help drive rapid resolution of critical technical issues.
This role offers the opportunity to work on cutting-edge server technologies that power large-scale AI and data center deployments. You will collaborate with world-class engineering teams, support internal AMD sites, and influence platform reliability through deep technical analysis, structured root cause investigations, and continuous process improvements. If you thrive in hands-on lab environments, enjoy solving difficult technical challenges, and want to make a direct impact on products deployed at scale, this is an exciting opportunity to be part of AMD's data center growth journey.
THE PERSONThe ideal candidate is a highly analytical problem solver who enjoys tackling complex technical challenges and driving issues to resolution. You are naturally curious, detail-oriented, and comfortable working through ambiguous problems involving hardware, firmware, and system interactions.
You excel in collaborative environments, effectively partnering with cross-functional teams while also taking ownership of independent investigations. Strong communication skills are essential, as you'll be expected to translate highly technical findings into clear recommendations and corrective actions for a broad range of stakeholders.
You are passionate about continuous learning, embrace emerging technologies and AI-assisted workflows, and enjoy helping others by creating documentation, sharing knowledge, and improving engineering processes.
KEY RESPONSIBILITIES- Within the first several months, become proficient with AMD server platforms, rack-level infrastructure, debug workflows, and failure analysis methodologies.
- Drive root cause investigations involving server, rack, and data center level failures across CPU, GPU, memory, PCIe, networking, storage, power, and thermal subsystems.
- Support internal AMD sites with system bring-up, issue triage, troubleshooting, and debug activities to accelerate deployment readiness.
- Reproduce field and manufacturing failures in laboratory environments and validate corrective actions that improve product quality and reliability.
- Partner with firmware, validation, design, manufacturing, and quality teams to drive cross-functional resolution of complex technical issues.
- Develop and maintain troubleshooting guides, standard operating procedures, and technical documentation that enable effective issue resolution across global teams.
- Contribute to Root Cause Analysis (RCA) and Failure Mode and Effects Analysis (FMEA) efforts, helping drive systemic improvements in reliability, serviceability, and testability.
- Leverage AI-assisted engineering tools and workflows to improve debug efficiency, accelerate log analysis, and enhance knowledge-sharing capabilities.
- Serve as a technical resource and mentor for internal teams by promoting best practices for failure isolation, structured debugging, and problem-solving.
PREFERRED EXPERIENCE- Hands-on experience debugging server platforms, rack-level systems, or data center infrastructure.
- Strong understanding of GPU server architectures and large-scale compute platforms.
- Experience troubleshooting CPU, GPU, memory, PCIe, networking, storage, power delivery, and thermal subsystems.
- Knowledge of platform bring-up procedures, manufacturing test environments, and hardware validation methodologies.
- Familiarity with BIOS/UEFI, BMC, IPMI, firmware debugging, and system management technologies.
- Experience using oscilloscopes, logic analyzers, protocol analyzers, power analyzers, and similar debug instrumentation.
- Ability to analyze system logs, firmware traces, and telemetry data to identify root causes and recommend corrective actions.
- Experience creating RCA, FMEA, troubleshooting, or technical documentation.
- Familiarity with AI-assisted engineering workflows, including log analysis, automated troubleshooting, and knowledge-base development.
- Strong collaboration, communication, and stakeholder management skills.
- Willingness to travel up to 25%.
ACADEMIC CREDENTIALS- Bachelor's or master's degree preferred in Electrical Engineering, Computer Engineering, Systems Engineering, or a related technical field.
LOCATIONSecaucus, NJ (Onsite)
THIS ROLE IS NOT ELIGIBLE FOR VISA SPONSORSHIP. #LI-CS1
Benefits offered are described: AMD benefits at a glance.