OpenAI

Data Center Hardware Quality & Reliability Engineer

OpenAI • $138K — $165K *
Telecommunications & Hardware
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • BS in engineering, physics, or equivalent; MS preferred
  • 8+ years in hardware quality/reliability or mission-critical infrastructure
  • Experience in managing field-failure and corrective action outcomes
  • Strong grasp of hardware architecture and various subsystems
  • Proficient in reliability statistics and analytical methods
  • Hands-on experience with FMEA/FTA and corrective-action processes
  • Skilled in SQL and Python/R for data analysis

Responsibilities

  • Own and build the data quality model for field telemetry and failures
  • Define key reliability metrics and their explicit metrics
  • Provide macro and micro level performance views to detect patterns
  • Lead triage and risk assessments of systemic field failures
  • Develop reliability and spare demand projections for hardware
  • Collaborate with MQE and NPI to enhance manufacturing test coverage
  • Initiate and verify upstream changes to mitigate recurrence

Benefits

  • Opportunity to influence design and operational changes
  • Mentorship potential as the team and responsibilities grow
  • Access to cutting-edge technologies in AI and data centers
  • Cross-functional collaboration with diverse teams
  • Engagement in high-impact reliability programs and initiatives
Full Job Description
About The Role

Own the end-to-end data-center hardware quality and reliability loop for OpenAI's 3P infrastructure and 1P current and next-gen platforms. Turn field failures into quantified risk, fast containment, verified root cause, improved MQE/NPI and manufacturing-test coverage, accurate spares forecasts, and upstream changes that prevent recurrence.

The first hire must combine practical hardware/system understanding, reliability engineering, data fluency, and cross-functional technical leadership at data-center scale.

Key Responsibilities
• Build and govern the field-quality data model across telemetry, tickets, RMA/repair, FA, firmware, configuration, supplier, and manufacturing genealogy.
• Define AFR, ASR, DPPM, MTBF/MTTR, repeat-repair, NTF, repair-cycle-time, and forecast-versus-actual metrics with explicit denominators and uncertainty.
• Provide fleet-level macro views and unit/FRU/cohort-level micro views; detect shifts and bound affected populations.
• Lead systemic field-failure triage, containment, failure analysis, 8D/CAPA, risk assessment, corrective-action verification, and recurrence monitoring.
• Develop cohort, life-data, reliability-growth, and spare-demand projections by product, FRU, supplier, configuration, geography, and age.
• Partner with MQE and NPI to convert field mechanisms into manufacturing-test coverage, screening/stress profiles, diagnostics, control plans, DFR/DFS requirements, FMEA/FTA, mission profiles, FRU strategy, and qualification gates.
• Close the loop by verifying whether upstream changes reduce field recurrence.
• Define supplier/CM FA standards, field-data contracts, scorecards, escalation paths, and closure evidence.
• Provide serviceability, TCO, and spares inputs without owning inventory execution or procurement.
• Create concise executive decision packages: population at risk, exposure, confidence, options, cost/risk, and recommendation.
• Run the cross-functional reliability council and, as the team grows, mentor the 1P and 3P Field Quality Engineers.

Qualifications
• BS in electrical, mechanical, computer, materials, reliability engineering, physics, or equivalent experience; MS preferred.
• 8+ years in hardware quality/reliability, server/rack systems, or mission-critical infrastructure; 3+ years owning field-failure, RMA, or CAPA outcomes.
• Solid working understanding of hardware and system architecture across board, tray, rack, firmware, telemetry, manufacturing test, and fleet behavior; deep expertise in every subsystem is not required.
• Reliability statistics: censored life data, Weibull/Poisson/binomial methods, confidence bounds, MTBF/MTTR, and reliability growth.
• Hands-on FMEA/FTA, accelerated or reliability-demonstration testing, 8D/CAPA, FA, and corrective-action verification.
• Working proficiency with SQL and Python/R or equivalent analytics tools.
• Ability to influence design, validation, operations, suppliers/CMs, and senior leaders without direct authority.

Preferred Skills
• GPU/AI server platforms, liquid cooling, high-power delivery, high-speed networking, rack integration, or data-center operations.
• Design for serviceability: FRU boundaries, diagnostics, repair workflows, tooling/access, and spares policy.
• Qualification-to-field correlation and mission-profile development.
• ODM/CM/supplier experience: FA quality, audit, QBR, and corrective-action governance.
• Linux/BMC/IPMI/Redfish logs and fleet telemetry.
• Leadership of a cross-generation reliability program or launch-readiness gate.

About OpenAI

OpenAI is an artificial intelligence research laboratory consisting of the for-profit corporation OpenAI LP and its parent company, the non-profit OpenAI Inc. The company was founded in 2015 by a group of technology leaders, including Elon Musk, Sam Altman, Greg Brockman, Ilya Sutskever, and John Schulman. OpenAI's mission is to develop and promote friendly AI for the betterment of humanity. The company has developed a number of cutting-edge AI technologies, including GPT-3, a language processing system that can generate human-like text. OpenAI has received funding from a number of high-profile investors, including LinkedIn co-founder Reid Hoffman and venture capitalist Peter Thiel.
Learn more about OpenAI
Size
100 employees
Industry
Founded
2015

Similar Jobs

More Jobs at OpenAI

More Telecommunications & Hardware Jobs

Find similar Data Center Hardware Quality & Reliability Engineer jobs: