OpenAI

AI Infrastructure Engineer, pAGI

OpenAI$150K — $180K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in software engineering or related field
  • Strong background in building or operating large-scale distributed systems
  • Experience with ML infrastructure or inference systems
  • Familiarity with GPU performance optimization
  • Self-motivated and capable of tackling open-ended problems

Responsibilities

  • Build and maintain infrastructure for large-scale training and evaluation
  • Develop shared inference and grading platforms with automated management features
  • Optimize compute scheduling to minimize idle GPU time
  • Diagnose and resolve bottlenecks in training and inference workflows
  • Create self-service tools and validation systems for researchers

Benefits

  • Collaborative work environment with researchers and engineering teams
  • Opportunity to impact AI technology and infrastructure development
  • Involvement in cutting-edge projects in the AI field
  • Potential for personal growth and development in AI systems engineering
  • Flexible working arrangements and a focus on work-life balance
Full Job Description
About the Team

pAGI Infra team builds and operates the systems that make large-scale model training and evaluation reliable, efficient, and easy to run. Our work spans distributed training infrastructure, inference and grading platforms, compute scheduling, and research tooling. We partner closely with researchers and engineering teams to turn new research needs into dependable infrastructure, improve GPU efficiency, and shorten the path from an experiment to a validated model.

About the Role

We're looking for an AI Systems Engineer to help scale the infrastructure behind our training and evaluation workflows. You'll own projects from identifying bottlenecks and designing solutions through deployment and operation. The work combines distributed systems engineering, performance optimization, and close collaboration with researchers. You might build a shared grading service, improve resource allocation across workloads, or bring a new training stack into production - directly improving how quickly and reliably research moves forward.

In this role, you will:
  • Build and operate infrastructure for large-scale training and evaluation, improving reliability, throughput, and resource efficiency.
  • Develop shared inference and grading platforms with automated capacity management, health monitoring, and visibility into performance.
  • Improve compute scheduling and resource allocation to reduce idle GPU time and help workloads recover quickly from failures.
  • Diagnose bottlenecks across training, inference, and orchestration, and work across teams to improve end-to-end performance.
  • Build self-service tools, automated validation, and observability that help researchers launch experiments, diagnose issues, and compare results with less manual intervention.

You might thrive in this role if you:
  • Are excited about the potential of personal AGI and want to build the infrastructure that enables it.
  • Have strong software engineering fundamentals and experience building or operating large-scale distributed systems.
  • Have experience in ML infrastructure, inference systems, GPU performance, or infrastructure tooling.
  • Are highly self-motivated and comfortable taking ownership of open-ended problems.
  • Enjoy debugging across system boundaries and using measurements to guide improvements in performance and reliability.


About OpenAI

OpenAI is an artificial intelligence research laboratory consisting of the for-profit corporation OpenAI LP and its parent company, the non-profit OpenAI Inc. The company was founded in 2015 by a group of technology leaders, including Elon Musk, Sam Altman, Greg Brockman, Ilya Sutskever, and John Schulman. OpenAI's mission is to develop and promote friendly AI for the betterment of humanity. The company has developed a number of cutting-edge AI technologies, including GPT-3, a language processing system that can generate human-like text. OpenAI has received funding from a number of high-profile investors, including LinkedIn co-founder Reid Hoffman and venture capitalist Peter Thiel.
Learn more about OpenAI
Size
100 employees
Industry
Founded
2015

Similar Jobs

More Jobs at OpenAI

More Enterprise Technology Jobs

Find similar AI Infrastructure Engineer, pAGI jobs: