OpenAI

Software Engineer, Infrastructure

OpenAI$130K — $155K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4+ years of relevant industry experience, including 2+ years leading complex projects or teams as an engineer or tech lead.
  • Strong software engineering skills in Python, Go, C++, Rust, or similar languages.
  • Experience in designing, operating, or scaling distributed systems or developer infrastructure.
  • Comfortable working in Linux environments and with Kubernetes, Terraform, CI/CD, and modern observability tools.
  • Excellent communication skills and adept at building consensus among diverse stakeholders.

Responsibilities

  • Design, build, and maintain reliable and high-performance systems across engineering.
  • Collaborate with engineers, product managers, and researchers to address evolving infrastructure needs.
  • Improve internal tooling and enhance the developer experience through automation.
  • Contribute to incident response, conduct postmortems, and develop best practices for system reliability.
  • Define technical strategy, architecture, and long-term goals with your team.

Benefits

  • Engage in high-impact projects that contribute to mission-critical services.
  • Work in a deeply collaborative environment with a variety of focus areas.
  • Opportunity to influence technical strategies and architecture at a leading technology organization.
  • Access to cutting-edge tools and technologies in distributed systems and cloud infrastructure.
Full Job Description
About the Team

We're hiring Software Engineers to join our broader Infrastructure organization, which supports multiple high-impact teams. Depending on your interests and experience, you could work on one of several focus areas-including Core Distributed Systems, Reliability Engineering, Observability, Developer Productivity or Cloud Infrastructure.

About the Role

All teams are deeply collaborative, work on mission-critical services, and are responsible for building distributed, scalable infrastructure to bring OpenAI's technology to the world through products like ChatGPT and the OpenAI API. You'll work closely with stakeholders to understand infrastructure, data and compute needs, setting the technical strategy that supports cutting-edge research and product development. This is a critical role for someone who is passionate about solving complex engineering problems at scale, ensuring their performance, scalability and reliability

Team Focus Areas
  • Distributed Systems: Owning and building important, highly scalable, available, performant, and reliable distributed systems (and their building blocks) to power the entire stack at OpenAI
  • Systems Engineering: Work across layers of the stack-debugging system bottlenecks, evolving core infrastructure, and solving novel problems in performance and scalability.
  • Reliability Engineering: Build scalable, fault-tolerant systems and lead efforts around service health, incident response, and resilience.
  • Observability: Design and maintain observability tooling (metrics, logs, tracing) to give teams visibility into production systems at scale.
  • Developer Productivity: Create tools, environments, and workflows that help engineers ship high-quality software faster and more safely.
  • Cloud Infrastructure: Own the cloud-native infrastructure (compute, networking, storage) that underpins all services and research workloads.
  • Databases: Building high performance, distributed database systems that power all of OpenAI's product stack.

In this role you will:
  • Design, build, and maintain reliable and performant systems used across engineering.
    Work with your team to define technical strategy, architecture, and long-term goals.
  • Collaborate with other engineers, product managers, and researchers to build infrastructure that meets evolving needs.
  • Improve internal tooling, automation, and developer experience.
  • Contribute to incident response, postmortems, and the development of best practices around system reliability and scalability.

You might thrive in this role if you:
  • Strong software engineering skills with experience in Python, Go, C++, Rust, or similar languages.
  • Experience designing, operating, or scaling distributed systems or developer infrastructure.
  • Comfort working in Linux environments, and with tools like Kubernetes, Terraform, CI/CD pipelines, and modern observability stacks.
  • Ability to navigate complex systems and a willingness to dig deep when debugging tricky issues.
  • Excellent communication and collaboration skills, especially in cross-functional settings.

Qualifications:
  • 4+ years of relevant industry experience, with 2+ years leading large scale, complex projects or teams as an engineer or tech lead
  • A passion for distributed systems at scale with a focus on reliability, scalability, security, and continuous improvement.
  • Excellent communication skills, with ability to build consensus among stakeholders both internally and externally.


About OpenAI

OpenAI is an artificial intelligence research laboratory consisting of the for-profit corporation OpenAI LP and its parent company, the non-profit OpenAI Inc. The company was founded in 2015 by a group of technology leaders, including Elon Musk, Sam Altman, Greg Brockman, Ilya Sutskever, and John Schulman. OpenAI's mission is to develop and promote friendly AI for the betterment of humanity. The company has developed a number of cutting-edge AI technologies, including GPT-3, a language processing system that can generate human-like text. OpenAI has received funding from a number of high-profile investors, including LinkedIn co-founder Reid Hoffman and venture capitalist Peter Thiel.
Learn more about OpenAI
Size
100 employees
Industry
Founded
2015

Similar Jobs

More Jobs at OpenAI

More Enterprise Technology Jobs

Find similar Software Engineer, Infrastructure jobs: