In this role, you'll set scientific direction and lead the development of techniques to improve speech-to-text accuracy for domain-specific use cases and optimize the performance and latency of end-to-end audio pipelines. You'll lead work across model evaluation and selection, fine-tuning, data and evaluation strategies, and real-time inference optimization. You'll partner closely with software engineers to turn scientific advances into production capabilities, while working with enterprise customers to understand how speech and audio systems perform in their environments. You'll turn improvements inspired by one customer's needs into capabilities that serve many, while influencing teams around a shared scientific vision.
We're a small, fast-moving team building AI-powered solutions at the intersection of Alexa AI capabilities and AWS cloud services. Our customers span healthcare, energy, retail, and insurance, and they're deploying in environments where the technical challenges are real and the feedback loops are immediate.
Key job responsibilities
- Set the scientific direction for improving speech-to-text accuracy across domain-specific use cases and real-world operating conditions
- Lead the evaluation, selection, adaptation, and fine-tuning of speech and audio models based on accuracy, latency, cost, reliability, and deployment constraints
- Drive improvements to the performance and latency of end-to-end audio pipelines, from signal processing and streaming through inference and transcription
- Define datasets, benchmarks, experiments, and evaluation methodologies that reflect customer use cases
- Lead the optimization of models and inference pipelines for reliable, real-time operation
- Work directly with enterprise customers to understand quality challenges, validate scientific improvements, and identify opportunities for broader platform capabilities
- Turn patterns from customer engagements into reusable models, evaluation methods, and scientific capabilities that scale across customers and industries
- Influence and collaborate with software engineering and partner teams to productionize models, measure their performance, and continuously improve deployed systems
- Write scientific and technical documents, lead reviews, and communicate experimental results and trade-offs to engineering, product, and business leaders
- Mentor Applied Scientists and engineers, participate in hiring, and raise the team's science and engineering standards
BASIC QUALIFICATIONS
- 3+ years of building machine learning models for business application experience
- PhD, or Master's degree and 6+ years of applied research experience
- Experience programming in Java, C++, Python or related language
- Experience with neural deep learning methods and machine learning
PREFERRED QUALIFICATIONS
- PhD or equivalent research experience, or a PhD and experience in patents or publications at top-tier peer-reviewed conferences or journals
- Experience leading a science team
- Experience applying machine learning to speech or audio processing, including generative models, speech enhancement, source separation, or related applications
- Familiarity with voice-processing techniques such as beamforming, acoustic echo cancellation, and noise reduction
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, CA, Sunnyvale - 192,200.00 - 260,000.00 USD annually
USA, WA, Seattle - 167,100.00 - 226,100.00 USD annually