Research Engineer - Data Infrastructure

ElevenLabs

$110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Experience building data-intensive systems, particularly for machine learning pipelines.
  • Strong skills in distributed data processing at scale, using technologies such as Kubernetes.
  • Ability to assess data quality and curation impact on model performance, and develop evaluation tools.
  • Demonstrated problem-solving skills through past projects or contributions to open-source.
  • Familiarity with data collection strategies and automated systems for data management.

Responsibilities

  • Own and optimize systems for data that supports frontier AI models.
  • Build and maintain large-scale data pipelines for data collection and transformation.
  • Train models within data processing pipelines to enhance efficiency and performance.
  • Design innovative data curation strategies to improve model quality.
  • Develop tools that enable researchers to work with extensive datasets efficiently.

Benefits

  • Fully remote role with global execution capability.
  • Option to work from offices in major cities: London, New York, San Francisco, and Warsaw.
  • Opportunity to influence and innovate in AI data infrastructure.
  • Collaborative environment focused on cutting-edge research.
  • Flexible work arrangements to support work-life balance.
Full Job Description
About the role

We are looking for a Research Engineer to join the research team at ElevenLabs, focused on the data infrastructure that powers our frontier AI models. The quality of our models is bounded by the quality and scale of the data behind them, and you will own the systems that make world-class data possible. You will thrive in this role if you enjoy:

  • Building large-scale data pipelines for collecting, processing, filtering, and transforming datasets used to train state-of-the-art models.
  • Training models used in our data processing pipelines, such as classifiers, quality filters, and labeling models.
  • Designing data curation strategies such as deduplication, quality scoring, labeling, and augmentation that measurably improve model performance.
  • Creating tooling and infrastructure that lets researchers explore and train on massive datasets quickly and reliably.


Requirements

We do not require any formal certifications or degrees. Instead, we are seeking enthusiastic engineers who can showcase solving impressively hard problems with artifacts such as past projects, designs, or GitHub contributions. Ideally, you bring:

  • Experience building data-intensive systems, ideally in support of machine learning training pipelines.
  • Strong engineering skills in distributed data processing at scale (e.g., Kubernetes, or custom pipelines over large datasets).
  • The capacity to autonomously evaluate how data quality, composition, and curation affect model outcomes, and to build the tooling to measure it.


Bonus: Experience building or operating web crawlers.

Location

This role is remote and can be executed globally. If you prefer, you can work from our offices in London, New York, San Francisco, and Warsaw.

#LI-Remote

Similar Jobs

More Jobs at ElevenLabs

More Information Technology Jobs

Find similar Research Engineer - Data Infrastructure jobs: