Google

System Hardware Reliability Engineer

Google$188K — $274K *
Technical Services
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Reliability Engineering, Data Science, Mechanical/Electrical Engineering, Applied Physics, or equivalent experience.
  • 8+ years in design for reliability techniques across consumer electronics.
  • 8+ years in hardware reliability engineering, including predictive analytics and physics of failure.
  • Master's degree preferred in relevant engineering or data science fields.
  • 7+ years experience in experimental design and execution of complex hardware products.
  • 6+ years statistical analysis experience, particularly in hardware reliability contexts.
  • Expertise in machine learning frameworks like PyTorch or TensorFlow for reliability forecasting.

Responsibilities

  • Design and implement advanced Prognostics and Health Management (PHM) algorithms using machine learning.
  • Develop stochastic degradation models to predict thermal impacts on hardware reliability.
  • Build frameworks for health state monitoring and anomaly detection with fleet telemetry data.
  • Collaborate with software teams to integrate predictive models into load management systems.
  • Lead reliability assessments and PoF modeling for various critical hardware components, including conducting FMEAs.

Benefits

  • Comprehensive health, dental, and vision coverage.
  • Retirement savings plan with company match.
  • Generous paid time off and vacation policy.
  • Professional development and educational assistance programs.
  • Work culture that promotes diversity and inclusion.
Full Job Description
Minimum qualifications:
  • Bachelor's degree in Reliability Engineering, Data Science, Mechanical/Electrical Engineering, Applied Physics, or equivalent practical experience.
  • 8 years of experience in applying Design for Reliability techniques, and working on multiple consumer electronics products.
  • 8 years of experience in hardware reliability engineering, physics of failure, and predictive analytics.

Preferred qualifications:
  • Master's degree in Reliability Engineering, Data Science, Mechanical/Electrical Engineering, Applied Physics, or equivalent practical experience.
  • 7 years of experience with experimental design and execution of complex hardware products.
  • 6 years of statistical analysis experience.
  • 6 years of experience working with manufacturing partners.
  • 5 years of experience with failure analysis techniques.
  • Expertise in physics-informed machine learning, Bayesian analysis, and utilizing frameworks like PyTorch or TensorFlow for reliability forecasting.


About the job

As a Reliability Engineer, you will play a key role in creating new consumer electronic products that meet a high bar for reliability and performance. You will work closely with the product management and design engineering teams to define standards, specify tests, and then supervise test execution and failure analysis. A broad engineering background and command of statistical methods will help to inform the design of new products. Your strong people management and communication skills will be key to ensuring adoption of your technical recommendations.

As a System Hardware Reliability Engineer, you will serve as the principal technical authority on hardware reliability, prognostics, and predictive analytics under dynamic thermal and environmental operating profiles. You will lead the development of sophisticated health monitoring models to evaluate the impact of elevated coolant temperatures, ambient air excursions, and dynamic workloads on the degradation and failure rates of compute accelerators, high-density servers, power electronics, and energy storage systems.

You will bridge classic Physics-of-Failure (PoF) modeling with machine learning to establish advanced Prognostics and Health Management (PHM) frameworks for our infrastructure. By developing algorithms that forecast remaining useful life and detect early-warning anomalies, you will perform system-level risk-benefit trade-offs between capacity efficiency and hardware lifespan. Your data-driven prognostic models will shape advanced cooling architectures and operational control strategies across our global computing footprint.

Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $188000 - $274000 (USD) 20% bonus target equity benefits

Learn more about benefits at Google .

Responsibilities
  • Design and implement Prognostics and Health Management (PHM) algorithms using physics-informed machine learning to forecast hardware degradation and Remaining Useful Life (RUL).
  • Develop stochastic degradation models to predict the impact of dynamic thermal and power envelopes on fleet reliability.
  • Build health state monitoring and anomaly detection frameworks leveraging massive fleet telemetry data to enable predictive maintenance.
  • Partner with software and controls teams to integrate predictive health models into automated load-management and thermal capping mechanisms.
  • Lead comprehensive reliability assessments and PoF modeling for silicon, interconnects, optics, thermal solutions and power delivery/battery systems. Drive Failure Modes and Effects Analyses (FMEAs) to identify vulnerabilities under extreme environmental operating conditions.


About Google

Google is a multinational technology company that specializes in Internet-related services and products. These include online advertising technologies, search engine, cloud computing, software, and hardware. Google was founded in 1998 by Larry Page and Sergey Brin while they were Ph.D. students at Stanford University. The company has grown tremendously since then and has become one of the most valuable companies in the world. Google's mission is to organize the world's information and make it universally accessible and useful.
Learn more about Google
Size
156,500 employees
Market Cap
$1,115.4 billion
Industry
Net Income
$40.2 billion
Founded
1998
5 Year Trend
+23.3%
Revenue
$182.5 billion
NASDAQ

Similar Jobs

More Jobs at Google

More Technical Services Jobs

Find similar System Hardware Reliability Engineer jobs: