Description
The AWS Neuron Science Core Algorithm team is looking for talented Applied Scientists to push the frontier of hardware-aware machine learning for Trainium and Inferentia, the AWS Machine Learning accelerators. In this rare role at the intersection of LLM modeling, large-scale training systems, and hardware/datatype co-design, you own model and algorithm decisions jointly with AWS custom silicon. You will own solutions end-to-end from research through production, publish at top venues, and work alongside distinguished engineers and scientists in a strategic growth area for AWS.
We actively work on these areas:
- Low-precision training and inference: MXFP8, MXFP4, and sub-4-bit training and inference recipes, stochastic rounding, and Trn4 datatype exploration.
- Trn-friendly architectures: model architectures that exploit hardware strengths without sacrificing quality.
- System-aware optimizers & efficient distributed systems: efficient optimizers and distributed systems that give the best accuracy, co-designed with the hardware.
- Foundation-model pre-training accuracy: end-to-end validation across model scales, catching training divergence early, and equivalence-checking tooling.
- GenAI for systems: RL post-training for NKI kernel generation, mitigating reward-hacking and accelerating under low precision on Trn.
Key job responsibilities
- Own scientific problems end-to-end - from research and experimentation through production impact - applying rigorous evaluation to complex, ill-defined problems at large scale.
- Develop production-quality code in PyTorch or JAX and integrate scientific components into large-scale training and inference systems with operational excellence and efficient resource usage.
- Partner with foundation-model, engineering, and hardware-architecture teams so your findings directly inform what gets built into Trainium and shipped in the product stack.
- Mentor fellow scientists and interns, give constructive peer reviews, and help shape team goals, priorities, and the technical roadmap.
- Author and publish research at top peer-reviewed venues (ICLR, NeurIPS, ICML, MLSys) and engage the broader scientific community.
BASIC QUALIFICATIONS
- PhD, or Master's degree and 6+ years of applied research experience
- Experience programming in Java, C++, Python or related language
- 3+ years of building machine learning models for business application experience
- Experience with neural deep learning methods and machine learning
PREFERRED QUALIFICATIONS
- Experience with modeling tools such as R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy etc.
- Experience with large scale distributed systems such as Hadoop, Spark etc.
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, CA, Cupertino - 192,200.00 - 260,000.00 USD annually