The OpportunityThe Speech AI Lab at Adobe Research is seeking a Senior Audio Research Scientist to join our speech generative AI and multimodal research efforts, based in San Francisco.
This role is ideal for researchers who stay deeply up to date with the latest AI model architectures and innovations and are motivated to push them forward in real-world creative systems. You will work alongside a world-class team of scientists and engineers to invent, evaluate, and scale next-generation audio and speech technologies-translating state-of-the-art research into products used by millions.
Our lab values independent research leadership, technical excellence, and impact. We actively support publication at top venues, collaboration with academia, and rapid iteration from research insight to deployed technology.
Responsibilities- Lead and execute independent research in speech generative AI, audio modeling, and multimodal learning (text and visuals), with a strong focus on modern large-scale model architectures.
- Design, analyze, and advance state-of-the-art AI models and training procedures, including foundation models, representation learning, and generative systems for speech and audio.
- Maintain strong technical currency with the latest developments in model architectures, optimization techniques, data scaling, and training strategies, and apply them creatively to audio and multimodal problems.
- Demonstrate deep understanding of both audio signal fundamentals and machine learning, enabling principled model design and insightful experimentation.
- Publish research at leading conferences and journals and pursue patent protection for impactful innovations.
- Build high-quality research prototypes that demonstrate technical rigor, scalability, and clear paths to product integration.
- Collaborate closely with engineering and product teams to ensure research outcomes translate into robust, real-world impact.
- Provide technical mentorship and help set a high bar for research quality and engineering excellence.
Requirements- Ph.D (preferred). or Master's degree in Computer Science, Electrical Engineering, or a related field, with a strong research focus in audio, speech, and machine learning.
- 3+ years of research experience in industry strongly preferred.
- Proven track record of independent research contributions, such as first-author publications or equivalent leadership in research projects.
- Deep expertise in audio and machine learning, including strong intuition for:
- Speech and audio generation
- Audio representations and modeling
- Training large-scale neural models
- Hands-on experience with modern AI architectures and training pipelines, and a demonstrated ability to quickly adopt and extend new techniques.
- Strong Python and deep learning development skills, with attention to performance, reproducibility, and experimental rigor.
- Excellent communication skills and the ability to clearly articulate complex technical ideas.
- A clear appetite for real-world impact, coupled with a commitment to technical excellence and high research standards.
Expected Pay Range:Our compensation reflects the cost of labor across several U.S. geographic markets, and we pay differently based on those defined markets. The U.S. pay range for this position is $142,700 -- $270,950 annually. Pay within this range varies by work location and may also depend on job-related knowledge, skills, and experience. Your recruiter can share more about the specific salary range for the job location during the hiring process.
In California, the pay range for this position is $187,100 - $270,950
At Adobe, for sales roles starting salaries are expressed as total target compensation (TTC = base + commission), and short-term incentives are in the form of sales commission plans. Non-sales roles starting salaries are expressed as base salary and short-term incentives are in the form of the Annual Incentive Plan (AIP).
In addition, certain roles may be eligible for long-term incentives in the form of a new hire equity award.