OpenAI

Software Engineer, Compute Foundations

OpenAI$150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in software engineering with a focus on distributed systems or infrastructure services
  • Proven experience in developing infrastructure systems utilizing Kubernetes APIs
  • In-depth understanding of hardware boot processes and management with knowledge of PXE, DHCP/DNS, and BMCs
  • Ability to design robust APIs and manage asynchronous workflows considering concurrency and failure
  • Strong diagnostic skills for troubleshooting performance across various systems and platforms
  • Excellent communication skills for cross-disciplinary collaboration

Responsibilities

  • Design and build Kubernetes-based controllers and distributed services to manage infrastructure
  • Define APIs and resource models for consistent lifecycle operations across hardware
  • Develop provisioning services for network boot, hardware management, and deployment
  • Implement lifecycle management for provisioning, upgrades, maintenance, and recovery
  • Create reliable systems for handling concurrent changes and recovering from failures
  • Enhance control-plane performance, API latency, and infrastructure state transition time
  • Integrate new hardware and sites into existing software platforms by collaborating with infrastructure teams

Benefits

  • Collaborative work environment with opportunities for cross-disciplinary projects
  • Exposure to cutting-edge GPU hardware and large-scale infrastructure challenges
  • Flexible work arrangements and a focus on work-life balance
  • Professional development opportunities, including training and conferences
  • Access to a culture that values innovation and deep technical expertise
Full Job Description
About the Role

You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong.

This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware.

We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack.

In this role, you will:
  • Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows.
  • Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers.
  • Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware, operating-system images, drivers, and host configuration.
  • Develop lifecycle management for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning, integrating with health and validation systems.
  • Design reliable reconciliation and recovery through concurrent changes, interrupted operations, and partial failures, with staged rollouts that limit disruption across nodes, racks, and clusters.
  • Improve control-plane throughput, API latency, and the time infrastructure takes to reach its desired state, while respecting the limits of site systems and provider APIs.
  • Build the software integrations that bring new sites and GPU hardware generations into the platform, partnering with hardware, networking, data-center, and other infrastructure teams.
You might thrive in this role if you:
  • Have strong software engineering fundamentals and experience designing, implementing, and owning production distributed systems or infrastructure services.
  • Have experience developing infrastructure systems that use Kubernetes APIs and reconciliation to manage resources.
  • Understand how a bare-metal node moves from power-on to a configured, workload-ready system, with depth in one or more areas such as PXE, DHCP/DNS, baseboard management controllers (BMCs), firmware, Linux, drivers, images, or configuration management.
  • Can design reliable APIs and asynchronous workflows, reasoning about concurrency, consistency, idempotency, and failures across service and provider boundaries.
  • Can diagnose reliability and performance problems across service, operating-system, and machine boundaries, turning production evidence into lasting software improvements.
  • Work effectively across engineering specialties and communicate system behavior and technical tradeoffs clearly.
Bonus points if you:
  • Have built infrastructure control planes that coordinate operations across multiple sites or regions.
  • Have worked with GPU or HPC infrastructure, including topology and shared dependencies across machines, racks, or clusters.
  • Have integrated multiple hardware platforms or infrastructure providers into a common service or resource model.

About OpenAI

OpenAI is an artificial intelligence research laboratory consisting of the for-profit corporation OpenAI LP and its parent company, the non-profit OpenAI Inc. The company was founded in 2015 by a group of technology leaders, including Elon Musk, Sam Altman, Greg Brockman, Ilya Sutskever, and John Schulman. OpenAI's mission is to develop and promote friendly AI for the betterment of humanity. The company has developed a number of cutting-edge AI technologies, including GPT-3, a language processing system that can generate human-like text. OpenAI has received funding from a number of high-profile investors, including LinkedIn co-founder Reid Hoffman and venture capitalist Peter Thiel.
Learn more about OpenAI
Size
100 employees
Industry
Founded
2015

Similar Jobs

More Jobs at OpenAI

More Information Technology Jobs

Find similar Software Engineer, Compute Foundations jobs: