Audio AI Engineer, #1085Multilingual Speech-to-Text Engineer - On-Device Model Optimization, #1085A Role with Purpose and ImpactThis role builds the speech recognition core of a mobile translation capability supporting a government agency's national security mission. The engineer will take large, high-quality speech-to-text models spanning many language families and adapt, compress, and optimize them so they run performantly on an iPhone - including handling the reality that speakers frequently mix in borrowed English terms mid-utterance, and the model needs to make a sound call on whether to transcribe those terms in English or in the source language's own transliteration.
This is an applied ML role, not a research-only position. The strongest candidate can move fluidly from raw audio data, to model adaptation and compression experiments, to a rigorous evaluation framework - and can clearly explain what they're building, why it's better than the status quo, and how they'll know it worked.
What This Role Is (and Isn't)This position owns the
speech-to-text model - its data, its training/adaptation, its size and latency on-device, and its accuracy across languages. It does not own iOS application development, translation (source-language-to-target-language), or the Swift/AVFoundation integration layer; those are handled by a separate mobile engineering function this role will collaborate closely with.
Key Responsibilities- Data pipelines: Ingest, clean, segment, label, and version multilingual audio and transcript data, with attention to code-switching and borrowed-word phenomena across the target language set.
- Model adaptation: Fine-tune and compress large ASR models (using LoRA/QLoRA, quantization, distillation, or other parameter-efficient and size-reduction techniques as appropriate) to fit iPhone-class memory, latency, and battery constraints, while preserving transcription quality.
- Dynamic, per-language deployment: Design model packaging so language-specific weights can be selected and downloaded on demand based on use-case context (e.g., an operator interviewing a Chinese speaker pulls only the Chinese ASR weights).
- Loanword/transliteration handling: Build and evaluate model behavior for deciding when a borrowed English term should be transcribed as-is versus rendered in the source language's transliteration or native equivalent.
- Evaluation: Build reproducible evaluation pipelines (word/character error rate, latency, robustness to accent/noise/speaking rate/code-switching) and clearly articulate results against defined success criteria for each language and deployment target.
- Documentation & communication: Produce clear model cards, dataset documentation, and evaluation write-ups that let technical and non-technical stakeholders understand what the model does, how it compares to alternatives, and what its risks and limitations are.
Required Qualifications- Bachelor's degree in Computer Science, Data Science, Machine Learning, Computational Linguistics, or a closely related field.
- Strong data-engineering background building production pipelines for large, messy, or unstructured audio/text datasets.
- Hands-on experience fine-tuning or adapting speech/audio models using parameter-efficient methods (LoRA, QLoRA, adapters) and/or model compression techniques (quantization, distillation, pruning) for constrained hardware.
- Practical experience with ASR/speech-to-text model development and evaluation across multiple languages, including error analysis under real-world conditions (accents, noise, code-switching).
- Strong Python and SQL skills; experience with PyTorch, Hugging Face Transformers/PEFT, torchaudio, librosa, or comparable tooling.
- Experience deploying and monitoring production ML systems, with an understanding of secure handling of sensitive audio, transcripts, and derived data in a regulated environment.
- Ability to clearly explain model behavior, tradeoffs, and limitations to both technical and non-technical stakeholders.
Preferred (Not Required)- Prior exposure to mobile/on-device ML deployment constraints (even without owning the mobile codebase directly).
- Experience with agentic or multi-step workflow orchestration involving model outputs, retrieval, or human review.
The estimated salary range for this position is
$80,000 - $160,000. This salary range is not a guarantee of compensation. The offered salary will be based on factors including relevant experience, geographic location, internal equity, and applicable contractual requirements. *Compensation may fall outside this range when appropriate.