Hardware Reliability Engineer

Meta

$138K — $165K *
Telecommunications & Hardware
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Electrical or Mechanical Engineering or related discipline
  • 6+ years in hardware reliability engineering
  • Proficient in reliability methodologies such as FMEA and HALT
  • Skilled in analyzing field failure data and root cause analysis
  • Experienced in collaborating with hardware suppliers for reliability evaluation
  • Effective communicator of complex technical findings

Responsibilities

  • Lead Design For Reliability (DFR) activities including DFMEA and derating
  • Develop reliability tests to identify design weaknesses in hardware
  • Establish Design Verification tests to evaluate environmental stresses
  • Oversee test execution with ODMs and recommend improvements based on past experiences
  • Translate test results into product life metrics and identify shortcomings
  • Utilize reliability statistics for risk decision making
  • Develop internal reliability test infrastructure and design experiments
  • Collaborate with cross-functional teams to mitigate design issues

Benefits

  • Opportunity to work on next-generation data center hardware
  • Influence hardware technology supporting Meta's AGI vision
  • Engage in high-impact projects connecting billions globally
  • Collaborative work environment with diverse engineering teams
  • Access to cutting-edge research and technology initiatives
Full Job Description
As a member of Meta Infrastructure's Hardware Product Integrity team you will work on next-generation data center hardware. You will be a part of futuristic projects including HW that will serve as the backbone for Meta's AGI vision, be in a position to influence HW technology that serves to connect billions of people across the world! In this role, you will drive reliability engineering efforts across server, storage, and networking hardware deployed in Meta's data centers, applying failure analysis, accelerated life testing, and reliability modeling to reduce field failures and improve hardware reliability.

Responsibilities

Lead DFR activities such as DFMEA, derating across various AI, compute and storage platforms
• Understanding technology that drives compute, storage, server hardware, and networking modules to develop reliability tests to bring out design weaknesses
• Establish Design Verification tests, to bring out environmental stress weaknesses in server design and ensure designs meet Meta's lifetime reliability metrics
• Work closely with ODMs to ensure and oversee tests are being executed as planned, suggest necessary improvements based on lessons learned from previous platforms
• Translate test results into meaningful product life metrics and highlight shortcomings in any metrics that are not met
• Utilize reliability statistics to help with decision making and quantifying risk and
• Lead the development of internal reliability test infrastructure to support initiatives and design of experiments
• Collaborate cross-functionally with Hardware Engineering, Release To Production, Thermal, and Failure Analysis teams to de-risk design issues

Minimum Qualifications
• Bachelor's degree in Electrical Engineering or Mechanical Engineering or a related discipline
• 6+ years of experience in hardware reliability engineering, including failure analysis and reliability testing of infrastructure hardware
• Experience applying reliability engineering methodologies such as FMEA, HALT, ALT, Weibull analysis, and MTBF modeling to infrastructure hardware
• Experience analyzing field failure data and translating findings into actionable root cause investigations and corrective actions
• Experience collaborating with hardware suppliers and contract manufacturers to evaluate component reliability and enforce qualification standards
• Experience communicating complex reliability findings and technical trade-offs to engineering and operations stakeholders through written reports and presentations

Preferred Qualifications
• Experience in silicon reliability and working on custom silicon is a plus
• MSc in Mechanical or Electrical Engineering or related disciplines
• Familiarity with data center environments is beneficial
• First-hand knowledge of server rack hardware is preferred

Similar Jobs

More Jobs at Meta

  • Business Engineer
    $120K — $145K *
    Seattle, WA 98115 (King County)
    Technical Services
    In-Person
  • Business Engineer
    $120K — $145K *
    New York, NY 10025 (New York County)
    Enterprise Technology
    In-Person
  • Facility Project Manager
    $100K — $120K *
    Rayville, LA 71269 (Richland County)
    Real Estate & Construction
    In-Person
  • Finance Manager, MSL FP&A
    $150K — $180K *
    Menlo Park, CA 94025 (San Mateo County)
    Finance & Insurance
    In-Person
  • Analytics Program Manager
    $150K — $180K *
    Menlo Park, CA 94025 (San Mateo County)
    Information Technology
    In-Person

More Telecommunications & Hardware Jobs

Find similar Hardware Reliability Engineer jobs: