Hardware Reliability Engineer

Meta

$130K — $155K *
Telecommunications & Hardware
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Electrical or Mechanical Engineering or related field
  • 6+ years in hardware reliability engineering, including testing and failure analysis
  • Proficient in reliability engineering methodologies (FMEA, HALT, ALT, Weibull analysis)
  • Experience translating field failure data into corrective actions
  • Strong collaboration skills with hardware suppliers and manufacturers
  • Ability to communicate complex reliability data effectively through reports and presentations

Responsibilities

  • Lead DFR activities including DFMEA and derating for various platforms
  • Develop reliability tests to identify weaknesses in server and storage hardware designs
  • Establish Design Verification tests to assess environmental stress on server designs
  • Oversee test execution with ODMs, suggesting improvements based on previous lessons
  • Translate test results into life metrics and identify unfulfilled metrics
  • Utilize reliability statistics for informed decision-making
  • Develop internal reliability test infrastructure to support design experiments

Benefits

  • Opportunities to work on groundbreaking technology linking billions globally
  • Collaboration with diverse teams in a fast-growing environment
  • Impact on Meta's AGI vision through hands-on hardware projects
  • Access to leading-edge test infrastructure and methodologies
  • Engagement with external ODMs to advance hardware reliability
Full Job Description
As a member of Meta Infrastructure's Hardware Product Integrity team you will work on next-generation data center hardware. You will be a part of futuristic projects including HW that will serve as the backbone for Meta's AGI vision, be in a position to influence HW technology that serves to connect billions of people across the world! In this role, you will drive reliability engineering efforts across server, storage, and networking hardware deployed in Meta's data centers, applying failure analysis, accelerated life testing, and reliability modeling to reduce field failures and improve hardware reliability.

Responsibilities

Lead DFR activities such as DFMEA, derating across various AI, compute and storage platforms
• Understanding technology that drives compute, storage, server hardware, and networking modules to develop reliability tests to bring out design weaknesses
• Establish Design Verification tests, to bring out environmental stress weaknesses in server design and ensure designs meet Meta's lifetime reliability metrics
• Work closely with ODMs to ensure and oversee tests are being executed as planned, suggest necessary improvements based on lessons learned from previous platforms
• Translate test results into meaningful product life metrics and highlight shortcomings in any metrics that are not met
• Utilize reliability statistics to help with decision making and quantifying risk and
• Lead the development of internal reliability test infrastructure to support initiatives and design of experiments
• Collaborate cross-functionally with Hardware Engineering, Release To Production, Thermal, and Failure Analysis teams to de-risk design issues

Minimum Qualifications
• Bachelor's degree in Electrical Engineering or Mechanical Engineering or a related discipline
• 6+ years of experience in hardware reliability engineering, including failure analysis and reliability testing of infrastructure hardware
• Experience applying reliability engineering methodologies such as FMEA, HALT, ALT, Weibull analysis, and MTBF modeling to infrastructure hardware
• Experience analyzing field failure data and translating findings into actionable root cause investigations and corrective actions
• Experience collaborating with hardware suppliers and contract manufacturers to evaluate component reliability and enforce qualification standards
• Experience communicating complex reliability findings and technical trade-offs to engineering and operations stakeholders through written reports and presentations

Preferred Qualifications
• Experience in silicon reliability and working on custom silicon is a plus
• MSc in Mechanical or Electrical Engineering or related disciplines
• Familiarity with data center environments is beneficial
• First-hand knowledge of server rack hardware is preferred

Similar Jobs

More Jobs at Meta

More Telecommunications & Hardware Jobs

Find similar Hardware Reliability Engineer jobs: