Research Staff, Voice AI Foundations

Deepgram

• $130K — $160K *
US-Anywhere
+ 2 other locationsRemote
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in machine learning or AI research, specifically related to audio processing and voice AI.
  • Strong background in statistical learning theory, particularly self-supervised and multimodal learning.
  • In-depth expertise in foundation model architectures and scaling techniques.
  • Experience designing data pipelines for processing large, diverse datasets.
  • Proven ability to conduct controlled experiments to evaluate architectural innovations.
  • Optimization expertise for real-world model deployment considering hardware constraints.
  • Familiarity with open-source contributions or research publications in speech/language AI.

Responsibilities

  • Pioneer new Latent Space Models (LSMs) for scalable voice AI solutions.
  • Develop next-generation neural audio codecs for efficient low bit-rate compression.
  • Create steerable generative models to synthesize diverse human speech.
  • Design embedding systems for interpretable latent space factorization.
  • Leverage latent recombination to generate synthetic audio at unprecedented scales.
  • Train multimodal speech-to-speech systems to understand diverse user inputs and provide empathetic responses.
  • Collaborate on robust model architectures and training schemes for billion-hour datasets.

Benefits

  • Collaborative and innovative work environment focused on ground-breaking AI research.
  • Opportunities for professional growth and skill development in AI and machine learning.
  • Contributions to high-impact projects that transform voice AI for users worldwide.
  • Access to cutting-edge resources and technologies in the field of audio AI.
Full Job Description
The Opportunity

Voice is the most natural modality for human interaction with machines. However, current sequence modeling paradigms based on jointly scaling model and data cannot deliver voice AI capable of universal human interaction. The challenges are rooted in fundamental data problems posed by audio: real-world audio data is scarce and enormously diverse, spanning a vast space of voices, speaking styles, and acoustic conditions. Even if billions of hours of audio were accessible, its inherent high dimensionality creates computational and storage costs that make training and deployment prohibitively expensive at world scale. We believe that entirely new paradigms for audio AI are needed to overcome these challenges and make voice interaction accessible to everyone.
The Role

As a Member of the Research Staff, you will pioneer the development of Latent Space Models (LSMs), a new approach that aims to solve the fundamental data, scale, and cost challenges associated with building robust, contextualized voice AI. Your research will focus on solving one or more of the following problems:

  • Build next-generation neural audio codecs that achieve extreme, low bit-rate compression and high fidelity reconstruction across a world-scale corpus of general audio.
  • Pioneer steerable generative models that can synthesize the full diversity of human speech from the codec latent representation, from casual conversation to highly emotional expression to complex multi-speaker scenarios with environmental noise and overlapping speech.
  • Develop embedding systems that cleanly factorize the codec latent space into interpretable dimensions of speaker, content, style, environment, and channel effects -- enabling precise control over each aspect and the ability to massively amplify an existing seed dataset through "latent recombination".
  • Leverage latent recombination to generate synthetic audio data at previously impossible scales, unlocking joint model and data scaling paradigms for audio. Endeavor to train multimodal speech-to-speech systems that can 1) understand any human irrespective of their demographics, state, or environment and 2) produce empathic, human-like responses that achieve conversational or task-oriented objectives.
  • Design model architectures, training schemes, and inference algorithms that are adapted for hardware at the bare metal enabling cost efficient training on billion-hour datasets and powering real-time inference for hundreds of millions of concurrent conversations.


The Challenge

We are seeking researchers who:
  • See "unsolved" problems as opportunities to pioneer entirely new approaches
  • Can identify the one critical experiment that will validate or kill an idea in days, not months
  • Have the vision to scale successful proofs-of-concept 100x
  • Are obsessed with using AI to automate and amplify your own impact

If you find yourself energized rather than daunted by these expectations-if you're already thinking about five ideas to try while reading this-you might be the researcher we need. This role demands obsession with the problems, creativity in approach, and relentless drive toward elegant, scalable solutions. The technical challenges are immense, but the potential impact is transformative.
It's Important to Us That You Have
  • Strong mathematical foundation in statistical learning theory, particularly in areas relevant to self-supervised and multimodal learning
  • Deep expertise in foundation model architectures, with an understanding of how to scale training across multiple modalities
  • Proven ability to bridge theory and practice-someone who can both derive novel mathematical formulations and implement them efficiently
  • Demonstrated ability to build data pipelines that can process and curate massive datasets while maintaining quality and diversity
  • Track record of designing controlled experiments that isolate the impact of architectural innovations and validate theoretical insights
  • Experience optimizing models for real-world deployment, including knowledge of hardware constraints and efficiency techniques
  • History of open-source contributions or research publications that have advanced the state of the art in speech/language AI

Similar Jobs

More Jobs at Deepgram

More Consumer Technology Jobs

Find similar Research Staff, Voice AI Foundations jobs: