OpenAI

Software Engineer, Fleet Infrastructure

OpenAI • $150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in cloud infrastructure or related field
  • Strong programming skills in relevant languages
  • Experience with hyperscale compute systems
  • Familiarity with Kubernetes for container orchestration
  • Hands-on experience in public cloud environments, especially Azure
  • User-focused mindset with an emphasis on execution
  • Understanding of AI/ML workloads is a plus

Responsibilities

  • Design and implement infrastructure components for GPU fleet operations
  • Build and maintain job scheduling and cluster management solutions
  • Develop snapshot delivery systems to enhance model startup times
  • Interface with research and product teams to gather workload requirements
  • Collaborate with hardware and infrastructure teams for high reliability services
  • Automate CI/CD processes for continuous integration and deployment systems
  • Support seamless delivery of research workflows through deployment systems

Benefits

  • Hybrid work model (3 days in office per week)
  • Relocation assistance for new employees
  • Opportunities to work with cutting-edge technology in AI/ML
  • Engagement in impactful projects that shape the future of AI
  • Collaboration with top-tier professionals in the field
Full Job Description
This role will support the fleet infrastructure team at OpenAI. The fleet team focuses on running the world's largest, most reliable, and frictionless GPU fleet to support OpenAI's general purpose model training and deployment. Work on this team ranges from
  • Maximizing GPUs doing useful work by building user-friendly scheduling and quota systems
  • Running a reliable and low maintenance platform by building push-button automation for kubernetes cluster provisioning and upgrades
  • Supporting research workflows with service frameworks and deployment systems
  • Ensuring fast model startup times though high performance snapshot delivery across blob storage down to hardware caching
  • Much more!

About the Role

As an engineer within Fleet infrastructure, you will design, write, deploy, and operate infrastructure systems for model deployment and training on one of the world's largest GPU fleet. The scale is immense, the timelines are tight, and the organization is moving fast; this is an opportunity to shape a critical system in support of OpenAI's mission to advance AI capabilities responsibly.

This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees.

In this role, you will:
  • Design, implement and operate components of our compute fleet including job scheduling, cluster management, snapshot delivery, and CI/CD systems.
  • Interface with researchers and product teams to understand workload requirements
  • Collaborate with hardware, infrastructure, and business teams to provide a high utilization and high reliability service


You might thrive in this role if you:
  • Have experience with hyperscale compute systems
  • Possess strong programming skills
  • Have experience working in public clouds (especially Azure)
  • Have experience working in Kubernetes
  • Execution focused mentality paired with a rigorous focus on user requirements
  • As a bonus, have an understanding of AI/ML workloads


About OpenAI

OpenAI is an artificial intelligence research laboratory consisting of the for-profit corporation OpenAI LP and its parent company, the non-profit OpenAI Inc. The company was founded in 2015 by a group of technology leaders, including Elon Musk, Sam Altman, Greg Brockman, Ilya Sutskever, and John Schulman. OpenAI's mission is to develop and promote friendly AI for the betterment of humanity. The company has developed a number of cutting-edge AI technologies, including GPT-3, a language processing system that can generate human-like text. OpenAI has received funding from a number of high-profile investors, including LinkedIn co-founder Reid Hoffman and venture capitalist Peter Thiel.
Learn more about OpenAI
Size
100 employees
Industry
Founded
2015

Similar Jobs

More Jobs at OpenAI

More Information Technology Jobs

Find similar Software Engineer, Fleet Infrastructure jobs: