Research Engineer - Inference

ElevenLabs

$110K — $130K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience deploying ML models in production, particularly for latency-sensitive applications.
  • Proficient in GPU programming with deep knowledge of optimization techniques (CUDA, Triton, TensorRT).
  • Ability to diagnose and resolve performance bottlenecks across multiple layers of the serving stack.
  • Strong experience with real-time data processing and streaming workloads.
  • Demonstrated ability through past projects or GitHub contributions.

Responsibilities

  • Deploy and optimize state-of-the-art AI models for real-time usage.
  • Enhance inference performance focusing on latency, throughput, and cost optimization.
  • Develop high-performance serving systems for critical streaming applications.
  • Create tooling to streamline the process of model deployment for researchers.
  • Manage the complete workflow from research to production implementation.

Benefits

  • Fully remote work option available globally.
  • Opportunity to work from major cities including London, New York, San Francisco, and Warsaw.
Full Job Description
About the role

We are looking for a Research Engineer to join the research team at ElevenLabs, focused on deploying and optimizing our frontier AI models in production. The quality of our models only matters if they can be served fast, reliably, and at scale. You will own the systems that turn research breakthroughs into real-time products used by millions. You will thrive in this role if you enjoy:
  • Deploying state-of-the-art models to production and owning the path from research checkpoint to serving infrastructure.
  • Optimizing inference performance across the stack, including latency, throughput, and cost, using techniques such as quantization, distillation, KV-cache optimization, batching strategies, and custom kernels.
  • Building and tuning high-performance serving systems for real-time, streaming workloads where every millisecond matters.
  • Creating tooling and infrastructure that lets researchers ship new models to production quickly, safely, and with confidence in their performance characteristics.


Requirements

We do not require any formal certifications or degrees. Instead, we are seeking enthusiastic engineers who can showcase solving impressively hard problems with artifacts such as past projects, designs, or GitHub contributions. Ideally, you bring:
  • Experience deploying and serving ML models in production, ideally for latency-sensitive or real-time applications.
  • Strong engineering skills in GPU programming and inference optimization (e.g., CUDA, Triton, TensorRT, or serving frameworks such as vLLM or SGLang).
  • The capacity to autonomously profile, diagnose, and eliminate bottlenecks across the serving stack, from model architecture to kernels to orchestration, and to build the tooling to measure it.


Location

This role is remote and can be executed globally. If you prefer, you can work from our offices in London, New York, San Francisco, and Warsaw.

#LI-Remote

Similar Jobs

More Jobs at ElevenLabs

More Enterprise Technology Jobs

Find similar Research Engineer - Inference jobs: