OpenAI

Software Engineer, Inference - Multi Modal

OpenAI • $150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in software engineering with a focus on inference systems
  • Strong understanding of GPU-based ML workloads
  • Familiarity with multimodal models and LLM inference
  • Experience with high-throughput, low-latency system optimization
  • Comfortable working in fast-paced, experimental environments
  • Knowledge of inference tooling like vLLM or TensorRT-LLM

Responsibilities

  • Design and implement infrastructure for multimodal model inference
  • Optimize image and audio input/output systems for performance
  • Transition experimental workflows to production-level services
  • Collaborate with researchers and product teams on deployment
  • Enhance system performance through GPU utilization and hardware abstraction

Benefits

  • Collaborative and cross-functional work environment
  • Opportunity to work on cutting-edge AI technologies
  • Involvement in innovative projects affecting real-time user interaction
  • Ability to influence and improve production-level services
  • Fast-paced, dynamic work culture fostering experimentation
Full Job Description
About the Role

We're looking for a software engineer to help us serve OpenAI's multimodal models at scale. You'll be part of a small team responsible for building reliable, high-performance infrastructure for serving real-time audio, image, and other MM workloads in production.

This work is inherently cross-functional: you'll collaborate directly with researchers training these models and with product teams defining new modalities of interaction. You'll build and optimize the systems that let users generate speech, understand images, and interact with models in ways far beyond text.

In this role, you will:
  • Design and implement inference infrastructure for large-scale multimodal models.
  • Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs.
  • Enable experimental research workflows to transition into reliable production services.
  • Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities.
  • Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers.

You might thrive in this role if you:
  • Have experience building and scaling inference systems for LLMs or multimodal models.
  • Have worked with GPU-based ML workloads and understand the performance dynamics of large models, especially with complex data like images or audio.
  • Enjoy experimental, fast-evolving work and collaborating closely with research.
  • Are comfortable dealing with systems that span networking, distributed compute, and high-throughput data handling.
  • Have familiarity with inference tooling like vLLM, TensorRT-LLM, or custom model parallel systems.
  • Own problems end-to-end and are excited to operate in ambiguous, fast-moving spaces.

Nice to Have:
  • Experience working with image generation or audio synthesis models in production.
  • Exposure to distributed ML training or system-efficient model design.

About OpenAI

OpenAI is an artificial intelligence research laboratory consisting of the for-profit corporation OpenAI LP and its parent company, the non-profit OpenAI Inc. The company was founded in 2015 by a group of technology leaders, including Elon Musk, Sam Altman, Greg Brockman, Ilya Sutskever, and John Schulman. OpenAI's mission is to develop and promote friendly AI for the betterment of humanity. The company has developed a number of cutting-edge AI technologies, including GPT-3, a language processing system that can generate human-like text. OpenAI has received funding from a number of high-profile investors, including LinkedIn co-founder Reid Hoffman and venture capitalist Peter Thiel.
Learn more about OpenAI
Size
100 employees
Industry
Founded
2015

Similar Jobs

More Jobs at OpenAI

More Information Technology Jobs

Find similar Software Engineer, Inference - Multi Modal jobs: