Join the AI Studios Engineering org within Prime Video and Amazon MGM Studios as a Machine Learning Engineer on CreativeFlux, the ML platform powering Nara, our AI-native content creation platform for professional animation and live action/VFX.
You will own the training and serving infrastructure that puts generative models in front of artists working on Prime Video's animated and live action productions. You'll work daily with Applied Scientists, animators, filmmakers, and storytellers, building the systems that turn new model architectures into production-grade tools artists can rely on under tight deadlines.
This role suits engineers who want their work to expand what storytelling can look like, putting new generative capabilities into the hands of the people building Prime Video's animated and live action shows.
Key job responsibilities
Build and operate ML training and serving infrastructure for the CreativeFlux platform. Strong hands-on experience with EKS / Kubernetes is preferred.
Design, deploy, and operate ML training and inference workloads on EKS, including GPU node groups, custom schedulers, autoscaling, resource quotas, and gang scheduling for multi-GPU jobs
Integrate generative ML models (image, video, audio) into production inference pipelines, partnering with Applied Scientists to take new architectures from research to scaled deployment
Build training and fine-tuning pipelines on EKS and SageMaker, including data preparation, distributed training, evaluation harnesses, and checkpointing strategies
Improve GPU utilization, throughput, and latency on the inference fleet through batching, quantization, model compilation, serving framework tuning, and Kubernetes-native scaling primitives (HPA, KEDA, custom controllers)
Build automated evaluation pipelines and quality metrics that catch model regressions before they reach Artists, including human-in-the-loop review where needed
Contribute to integrations with Digital Content Creation (DCC) tools such as Adobe Animate, After Effects, Maya, and Storyboard Pro at the points where models are surfaced to creators
Own operational health of model endpoints on EKS, including monitoring, alerting, on-call response, capacity planning, and post-incident analysis
Write clean, well-tested code and participate in design and code reviews
Stay current with generative ML, distributed training, inference optimization, and GPU efficiency techniques, and apply them where they move team metrics
A day in the life
As an MLE on the AI Studios team, you'll spend most of your time deploying and tuning generative ML models, optimizing inference pipelines, and building the EKS infrastructure that lets Applied Scientists move new architectures into production. A typical day might include scaling a video generation endpoint to handle a production deadline, debugging a multi-GPU training job that's checkpointing too slowly, or working with a Scientist to take a new image model from a research notebook to a serving fleet. You'll work with other Nara platform teams who owns other building blocks, making sure ML capabilities land cleanly in the workflows artists actually use. You'll participate in team standups, model review sessions, on-call rotations, and platform-wide design reviews with the other Nara teams. You'll learn from experienced engineers and scientists and gradually take on more complex challenges as you grow.
BASIC QUALIFICATIONS
- 3+ years of non-internship professional software development experience
- 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
- 1+ years of software development engineer or related occupational experience
- 1+ years of designing and developing large-scale, multi-tiered, multi-threaded, embedded or distributed software applications, tools, systems, and services using: C#, C++, Java, or Perl experience
- 1+ years of Object Oriented Design experience
- Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field
- Experience programming with at least one software programming language
PREFERRED QUALIFICATIONS
- 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
- Bachelor's degree in computer science or equivalent
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, CA, Culver City - 143,700.00 - 194,400.00 USD annually
USA, NY, New York - 158,100.00 - 213,800.00 USD annually