OpenAI

Software Engineer, Model Runtime

OpenAI$150K — $180K *
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in systems programming with proficiency in C++, Rust, or Python.
  • Experience building or optimizing system runtimes and distributed systems.
  • Deep understanding of modern large language model (LLM) inference mechanisms.
  • Ability to quantitatively analyze latency, throughput, and memory utilization.
  • Proficiency in profiling and debugging across various hardware-software layers.
  • Capability to create clean abstractions while optimizing for performance on special hardware.
  • Strong collaboration skills to drive complex technical solutions across teams.

Responsibilities

  • Design and implement the LLM inference runtime for cutting-edge models on custom silicon.
  • Develop high-performance scheduling, memory management, and execution orchestration.
  • Create distributed execution strategies, ensuring effective communication and synchronization across resources.
  • Optimize performance metrics, including latency and hardware utilization across varied workloads.
  • Collaborate with cross-functional teams to co-design performance-enhancing interfaces.
  • Facilitate the introduction of novel features and capabilities in a reliable production runtime.
  • Implement tools for profiling and benchmarking runtime performance.
  • Debug and resolve intricate performance and reliability challenges across the software-hardware spectrum.

Benefits

  • Opportunity to work on innovative projects with cutting-edge technology.
  • Collaborative work environment with cross-functional team interactions.
  • Focus on production quality and operational excellence.
  • Direct impact on shaping the performance of new model capabilities.
  • Potential to contribute to advancements in AI and silicon technology.
Full Job Description
About the Role

You will build the model runtime within the inference engine that executes complex, frontier models at scale on OpenAI's custom silicon. The runtime will sit between models running on the hardware and the upper layers of the cluster serving software stack, translating demanding inference workloads into efficient execution while optimizing for throughput, latency, utilization, and reliability.

You will work across model architecture, distributed systems, compilers, kernels, and silicon to design a production-grade runtime comparable in ambition to systems such as vLLM and SGLang, but customized and optimized for OpenAI's AI accelerator. Your work will shape how new model capabilities map onto the platform and how quickly custom silicon can deliver meaningful performance in production.

In this role, you will:
  • Design and implement the LLM inference runtime for frontier models running on custom silicon.
  • Build scheduling, continuous batching, memory management, KV-cache management, and execution orchestration for high-performance inference.
  • Develop distributed execution strategies across chips, hosts, and racks, including model partitioning, communication, and synchronization.
  • Optimize end-to-end latency, throughput, memory efficiency, and hardware utilization across diverse model architectures and serving workloads.
  • Partner with kernel, compiler, architecture, and silicon teams to co-design interfaces and remove performance bottlenecks across the stack.
  • Enable new model features, execution patterns, numerical formats, and hardware capabilities in a reliable production runtime.
  • Create profiling, observability, benchmarking, and performance-modeling tools that make runtime behavior measurable and actionable.
  • Debug complex correctness, performance, and reliability issues spanning model code, runtime software, communication layers, and hardware.
  • Turn workload insights into clear requirements for future generations of silicon and system architecture.

You might thrive in this role if:
  • Have strong systems programming experience in C++, Rust, Python, or comparable performance-oriented environments.
  • Have built or optimized runtimes, distributed systems, compilers, kernels, model-serving infrastructure, or adjacent systems software.
  • Understand modern LLM inference, including prefill and decode behavior, batching, KV-cache tradeoffs, and model parallelism.
  • Can reason quantitatively about latency, throughput, compute intensity, memory bandwidth, communication, and utilization.
  • Are comfortable profiling and debugging performance across multiple layers of a hardware-software stack.
  • Can design clean abstractions while retaining the low-level control needed to extract performance from specialized hardware.
  • Work effectively across model, systems, compiler, kernel, and hardware teams to drive ambiguous technical problems to closure.
  • Care about production quality, including correctness, observability, reliability, maintainability, and graceful behavior at scale.

To comply with U.S. export control laws and regulations, candidates for this role may need to meet certain legal status requirements as provided in those laws and regulations.

About OpenAI

OpenAI is an artificial intelligence research laboratory consisting of the for-profit corporation OpenAI LP and its parent company, the non-profit OpenAI Inc. The company was founded in 2015 by a group of technology leaders, including Elon Musk, Sam Altman, Greg Brockman, Ilya Sutskever, and John Schulman. OpenAI's mission is to develop and promote friendly AI for the betterment of humanity. The company has developed a number of cutting-edge AI technologies, including GPT-3, a language processing system that can generate human-like text. OpenAI has received funding from a number of high-profile investors, including LinkedIn co-founder Reid Hoffman and venture capitalist Peter Thiel.
Learn more about OpenAI
Size
100 employees
Industry
Founded
2015

Similar Jobs

More Jobs at OpenAI

More Consumer Technology Jobs

Find similar Software Engineer, Model Runtime jobs: