Amazon's Artificial General Intelligence (AGI) organization is seeking an Applied Scientist III to advance the science of Responsible AI evaluation for large language models and generative AI. In this role, you will lead the design and development of rigorous evaluation methods, benchmarks, and metrics that measure the safety, fairness, robustness, and trustworthiness of frontier models. You will work with large-scale datasets, modern deep learning frameworks, and world-class scientists and engineers to turn research into evaluation systems that shape model launch decisions at Amazon scale.
Key job responsibilities
- Lead the design and implementation of evaluation frameworks, benchmarks, and metrics for responsible AI, including safety, fairness, robustness, and harmful content.
- Build scalable automated evaluation pipelines for large language models, including model-based and human-in-the-loop evaluation.
- Partner with pretraining, post-training, and product teams to translate evaluation results into model improvements and launch decisions.
- Conduct rigorous experimentation and statistical analysis, and publish research at top venues.
- Mentor junior scientists and help raise the scientific bar of the team.
- Champion responsible AI practices across the model development lifecycle.
About the team
The AGI Responsible AI (RAI) team builds the science and systems that make Amazon's large language models safe, fair, and trustworthy. We work on problems spanning safety evaluation, content moderation, watermarking, bias mitigation, and alignment. Our team values scientific rigor, customer obsession, and rapid iteration, and we collaborate closely with pretraining, post-training, and product teams across AGI.
BASIC QUALIFICATIONS
- 3+ years of building machine learning models for business application experience
- PhD, or Master's degree and 6+ years of applied research experience
- Experience programming in Java, C++, Python or related language
- Experience with neural deep learning methods and machine learning
PREFERRED QUALIFICATIONS
- Experience with modeling tools such as R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy etc.
- Experience with large scale distributed systems such as Hadoop, Spark etc.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, CA, Sunnyvale - 192,200.00 - 260,000.00 USD annually
USA, MA, Boston - 167,100.00 - 226,100.00 USD annually