OpenAI

Software Engineer, Trainium

OpenAI$150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of relevant engineering experience in ML systems, compilers, kernels, or performance engineering.
  • Strong systems programming skills with a focus on performance-critical software.
  • Experience with GPU, TPU, Trainium, or other specialized accelerator architectures.
  • Ability to analyze performance issues across the software stack, from hardware to ML frameworks.
  • Demonstrated ownership of complex, ambiguous technical problems end-to-end.
  • Familiarity with AWS Trainium or the AWS Neuron SDK is a plus.
  • Experience with contributions to ML frameworks, compilers, or kernel development for accelerators is advantageous.

Responsibilities

  • Build and optimize the inference stack for AWS Trainium.
  • Develop high-performance kernels for critical model operations.
  • Enhance compiler support for efficient target execution on Trainium.
  • Create systems for executing and optimizing the model forward pass on Trainium.
  • Profile workloads and identify performance bottlenecks across the tech stack.
  • Collaborate with teams to integrate new models and architectures into Trainium.
  • Work across hardware/software boundaries to enhance performance from AI accelerators.
  • Manage complex performance and systems problems from investigation to production.

Benefits

  • Opportunity to work at the forefront of AI and machine learning technology.
  • Collaborative and innovative work environment with cross-functional teams.
  • Chance to tackle challenging performance issues in large-scale AI systems.
  • Access to cutting-edge tools and technologies in performance engineering.
  • Potential for influence over the development and integration of next-generation AI models.
Full Job Description
About the Role

As a Software Engineer, Trainium, you will help bring OpenAI's inference workloads to AWS Trainium and build the software stack required to run cutting-edge frontier models efficiently on the platform.

This is a deeply technical, cross-stack role spanning kernels, compilers, and model execution. You will work on the systems needed to support OpenAI's inference stack on Trainium, including developing and optimizing high-performance kernels, improving compiler support, and enabling efficient execution of the model forward pass.

You'll work closely with engineers across inference, compilers, kernels, and ML systems to identify performance bottlenecks and build the software needed to take full advantage of Trainium. The work may range from low-level hardware-specific optimization to compiler and runtime improvements to integrating new model architectures into the inference stack.

If you enjoy working at the intersection of ML systems, compilers, kernels, and accelerator hardware, this role is for you. We're looking for engineers who are self-directed, comfortable operating across abstraction layers, and excited to solve challenging performance problems for frontier-scale AI systems.

In This Role, You Will
  • Build and optimize OpenAI's inference stack for AWS Trainium.
  • Develop high-performance kernels for critical model operations and workloads.
  • Extend and improve compiler support to efficiently target Trainium hardware.
  • Build the systems necessary to execute and optimize the model forward pass on Trainium.
  • Profile workloads and identify bottlenecks across kernels, compiler-generated code, runtime, and model execution.
  • Partner with inference and ML systems teams to bring new models and architectures onto Trainium.
  • Work across the hardware/software boundary to unlock performance and capabilities from specialized AI accelerators.
  • Own complex performance and systems problems end-to-end, from investigation through production deployment.

We're Looking for a Track Record Of
  • 3+ years of relevant engineering experience, ideally in ML systems, compilers, kernels, runtimes, or performance engineering.
  • Strong systems programming fundamentals and experience writing performance-critical software.
  • Experience working with GPU, TPU, Trainium, or other specialized accelerator architectures.
  • Ability to reason about performance across multiple layers of the stack, from hardware and kernels through compilers and ML frameworks.
  • Owning technically ambiguous problems end-to-end and learning new hardware and software domains as needed.
  • Bonus: experience with AWS Trainium or the AWS Neuron SDK.
  • Bonus: contributions to ML frameworks such as PyTorch or JAX, compiler infrastructure such as LLVM, MLIR, XLA, or Triton, or experience developing kernels for specialized accelerators.

About OpenAI

OpenAI is an artificial intelligence research laboratory consisting of the for-profit corporation OpenAI LP and its parent company, the non-profit OpenAI Inc. The company was founded in 2015 by a group of technology leaders, including Elon Musk, Sam Altman, Greg Brockman, Ilya Sutskever, and John Schulman. OpenAI's mission is to develop and promote friendly AI for the betterment of humanity. The company has developed a number of cutting-edge AI technologies, including GPT-3, a language processing system that can generate human-like text. OpenAI has received funding from a number of high-profile investors, including LinkedIn co-founder Reid Hoffman and venture capitalist Peter Thiel.
Learn more about OpenAI
Size
100 employees
Industry
Founded
2015

Similar Jobs

More Jobs at OpenAI

More Information Technology Jobs

Find similar Software Engineer, Trainium jobs: