About the role
We are seeking a highly skilled Staff Failure Analysis Engineer to lead complex investigations of server and datacenter hardware failures, driving root cause identification and corrective actions that improve product quality and reliability. In this role, you will work hands-on with server systems and components including CPUs, GPUs, DIMMs, NVMe drives, FPGA, NICs, and power supplies, performing troubleshooting, reliability testing, data analysis, and debug activities throughout the product lifecycle. You will partner closely with engineering, manufacturing, quality, and test teams to resolve issues, enhance manufacturing processes, support new product introductions, and provide technical leadership in a fast-paced, high-volume manufacturing environment. The ideal candidate brings strong server hardware expertise, failure analysis experience, and a passion for delivering reliable, high-performance technology solutions
What you will do
- Lead complex FA investigations from failure reproduction through root-cause identification and corrective action.
- Analyze test logs, MRS/RMA records, repair history, and manufacturing test data.
- Perform hardware-level troubleshooting of server systems, PCBA, CPU, DIMM, GPU/EAM, PCIe, NVMe, FPGA, NIC, PSU, and related components.
- Determine whether failures are related to hardware, firmware, software, workmanship, process, or design/CRD changes.
- Partner with PE/MQE/QC/TE/MFG to drive 4D investigations containment and permanent corrective actions.
- Review repair activities and ensure all failure and repair history is accurately captured.
- Identify recurring issues and drive improvements to test coverage and manufacturing processes.
- Support new product introduction and production ramp activities.
- Provide technical leadership and mentoring to FA/test engineers and manufacturing teams.
- Execute server hardware reliability system testing, reliability stresses, failure analysis, and statistical analysis through all phases of the product life cycle working with cross-functional teams, including hardware developers and system engineers.
- Hands-on Hardware reliability system testing, reliability stresses, failure analysis, debug, and statistical analysis.
- Work with engineering and other cross-functional team management to define operation project requirements, solutions, and schedules.
- Develop innovative techniques/approaches to accelerate failure identification and mechanism understanding and support technology transfer to high volume manufacturing.
- Participate in product design and reliability reviews during new product development to ensure the robustness of product design and manufacturing processes.
What You Bring
- In-depth of servers and network technologies, know-how and hands-on debug skills for datacenter servers.
- Strong troubleshooting, communication, and cross-functional leadership skills.
- Understanding of server component installation/uninstallation, connection, and networking setup.
- Linux ,Bash script, Windows power shell, and Python knowledge are strongly preferred.
- Knowledge of test methodologies, writing test plans, creating test cases, and debugging.
- Experience working in a high-volume manufacturing environment is preferred.
- Strong Hands-on experience in hardware/system repair debug and failure analysis.
- Experience analyzing statistical quality and reliability data.
- Demonstrated ability and experience in writing reports, business correspondence, and procedure manuals.
- Experience communicating with third party vendors while maintaining and developing relationships both internal and external.
- Strong interpersonal skills and the ability to interact with people at various levels in the organization.
- Ability to work in a fast-paced environment.
- Good analytical hands-on skill.
- Collaborative, flexible, and adaptable.
- Bachelor’s degree in electrical engineering, Computer Engineering, Systems Engineering or an equivalent combination of professional experience.
- 2+ years of experience executing tests and automating test cases.
- 5+ years of computer hardware relevant industry experience.
ZT Systems, a Sanmina Company, assesses market data to ensure a competitive compensation package is created for all our employees. The typical base salary for this position is expected to be between $93,000 and $136,400annually. If hired, the final base salary will be determined on an individual basis taking into consideration experience, skills, knowledge, education and/or certifications.
Base salary is just one component of ZT Systems total rewards philosophy. We take pride in offering a wide range of benefits and perks that appeal to the variety of needs across our diverse employee base. Other rewards may include bonus, paid time off, 401(k) retirement savings plan, tuition reimbursement, wellbeing resources, and more.
We are dedicated to building a diverse, inclusive, and authentic workplace, so if you’re excited about this opportunity but your experience doesn’t align perfectly with every qualification in the job description, we encourage you to apply anyway. You may be just the right candidate for this or other roles.
#LI-DH1