OpenAI

Software Engineer, Hardware Health

OpenAI • $150K — $180K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years of industry experience in software or infrastructure engineering.
  • Strong proficiency with Python and shell scripting.
  • Experience building large-scale distributed systems or infrastructure platforms.
  • Comfort with operational data exploration using SQL, PromQL, or similar tools.
  • Ability to create reproducible analyses and operational tooling.
  • Strong systems debugging instincts and ownership mindset.

Responsibilities

  • Define and maintain health signals for GPUs, CPUs, networking, and platform infrastructure.
  • Build and evolve health checks that detect, remediate, and verify failures at scale.
  • Ensure critical health checks execute with minimal latency to maximize workload uptime.
  • Investigate hardware failures and system-level issues in large-scale compute environments.
  • Own node lifecycle workflows including drain, quarantine, repair, and return-to-service processes.
  • Create automation and tooling for global cluster management with minimal manual intervention.
  • Collaborate with workload, reliability, and provider teams to integrate health signals into training and inference systems.

Benefits

  • Opportunity to work on cutting-edge technology within a high-impact team.
  • Potential for career growth in a fast-evolving AI environment.
  • Access to advanced tooling and infrastructure for problem-solving.
  • Collaborative atmosphere with diverse teams across the organization.
  • Focus on continuous learning and improvement within the role.
Full Job Description
About the Team

The Hardware Health and Observability team owns the end-to-end health lifecycle of OpenAI's global compute fleet.

Our mission is to maximize healthy, usable compute across accelerator vendors, generations, cloud providers, and regions through reliable health signals, automated remediation, and scalable operational tooling.

We build the systems that observe, detect, remediate, and verify hardware issues across GPUs, CPUs, networking, and platform infrastructure, enabling frontier model training and inference workloads to run reliably at hyperscale. We are the last line of defense for the success of OAI's production and research workloads.

About the Role

On the Hardware Health and Observability team, you'll build critical infrastructure that keeps OpenAI's largest compute clusters healthy and operational at scale.

Even small numbers of unhealthy systems can impact large-scale training and inference workloads. This team focuses on minimizing downtime, improving fleet efficiency, and ensuring compute resources remain continuously available to researchers and product teams.

Engineers on this team own problems end-to-end, from defining health signals and debugging failures to building automated remediation systems that operate across millions of GPUs globally.

In this role, you will:
  • Define and maintain health signals across GPUs, CPUs, networking, and platform infrastructure.
  • Build and evolve health checks that detect, remediate, and verify failures at scale.
  • Ensure critical health checks execute with minimal latency to maximize workload uptime.
  • Investigate hardware failures and system-level issues across large-scale compute environments.
  • Own node lifecycle workflows including drain, quarantine, repair, RMA, and return-to-service processes.
  • Build automation and tooling that enables global cluster management with minimal manual intervention.
  • Partner with workload, reliability, and provider teams to integrate health signals into training and inference systems.
You might thrive in this role if you have:
  • 7+ years of industry experience in software or infrastructure engineering.
  • Strong proficiency with Python and shell scripting.
  • Experience building large-scale distributed systems or infrastructure platforms.
  • Comfort digging into noisy operational data using SQL, PromQL, or similar tooling.
  • Experience building reproducible analyses and operational tooling.
  • Strong systems debugging and operational instincts with an ownership mindset.
Bonus if you have:
  • Experience with low-level hardware systems and Linux tooling (e.g. PCIe, InfiniBand, RoCE, networking, power management, kernel performance tuning, FW/SW debugging).
  • Experience operating or debugging large-scale GPU or accelerator clusters.
  • Expertise in network operations, observability, or systems telemetry.
  • Experience with automated remediation systems or fleet lifecycle management.
  • Experience improving reliability, utilization, or workload uptime in distributed compute environments.


About OpenAI

OpenAI is an artificial intelligence research laboratory consisting of the for-profit corporation OpenAI LP and its parent company, the non-profit OpenAI Inc. The company was founded in 2015 by a group of technology leaders, including Elon Musk, Sam Altman, Greg Brockman, Ilya Sutskever, and John Schulman. OpenAI's mission is to develop and promote friendly AI for the betterment of humanity. The company has developed a number of cutting-edge AI technologies, including GPT-3, a language processing system that can generate human-like text. OpenAI has received funding from a number of high-profile investors, including LinkedIn co-founder Reid Hoffman and venture capitalist Peter Thiel.
Learn more about OpenAI
Size
100 employees
Industry
Founded
2015

Similar Jobs

More Jobs at OpenAI

More Information Technology Jobs

Find similar Software Engineer, Hardware Health jobs: