Staff ML Engineer - AWS Trainium & SageMaker

Robots and Pencils

$150K — $180K *
US-AnywhereRemote in Canada
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Strong hands-on PyTorch experience, ideally with distributed or multi-device training
  • Production experience with Amazon SageMaker for training and/or inference
  • Proficiency in debugging near the hardware layer, focusing on device-specific compilation
  • Familiarity with AWS Trainium or Inferentia, with a track record of adapting to new hardware targets
  • Solid Python programming skills and experience in a client-facing, production environment

Responsibilities

  • Train and operate models on Amazon SageMaker utilizing AWS Trainium
  • Write and optimize PyTorch training code with an understanding of its execution on Trainium
  • Diagnose training issues related to hardware, distinguishing them from data or code problems
  • Translate training requests into effective, cost-aware training pipelines
  • Tune distributed training for better throughput and cost efficiency on SageMaker
  • Collaborate with client and internal teams to deliver production training workloads

Benefits

  • Opportunity to work with cutting-edge AWS Trainium hardware
  • Engagement in complex, production-level machine learning projects
  • Collaboration with internal engineering and client teams
  • Possibility to deepen knowledge of custom silicon and distributed training
  • A dynamic environment focused on real-world impact rather than just experimentation
Full Job Description
The Role
We're looking for an engineer who can operate and train models on Amazon SageMaker running on AWS Trainium, AWS's custom silicon built specifically for large-scale model training. This isn't a role where you call an API and wait. You'll be walking up the stack: understanding what a training request actually looks like at the Trainium hardware and compiler level, then carrying that understanding all the way up through PyTorch training code and into a production SageMaker pipeline.
PyTorch is the backbone of this work. If you know the framework deeply and you're comfortable reasoning about how your code actually behaves on custom accelerator hardware rather than treating it as a black box, this role is built around that skill set specifically.

What You'll Do
  • Train and operate models on Amazon SageMaker with AWS Trainium as the underlying compute
  • Write and optimize PyTorch training code with a real understanding of how it compiles and executes on Trainium (NeuronCore architecture, compiler behavior, memory and throughput tradeoffs)
  • Diagnose training run issues that show up specifically because of the hardware, not just the model, distinguishing a data or code problem from a compiler or device-level one
  • Translate a request for "a Trainium job" into an actual working, cost-aware training pipeline, end to end
  • Tune distributed training runs for throughput and cost on SageMaker's training infrastructure
  • Work directly with client and internal engineering teams to scope and deliver real production training workloads, not experiments that stay in a notebook

What You'll Bring
  • Strong, hands-on PyTorch experience, ideally including distributed or multi-device training
  • Production experience with Amazon SageMaker for training and/or inference
  • Comfort working close to the hardware layer: you understand device-specific compilation and can debug issues that are actually about the accelerator, not just the model
  • AWS Trainium or Inferentia (Neuron SDK) experience is a strong plus; if you don't have it yet but have deep PyTorch and a track record of picking up new hardware targets fast, we want to talk to you
  • Solid Python fundamentals and comfort operating in a client-facing, production engineering environment

Similar Jobs

More Jobs at Robots and Pencils

More Information Technology Jobs

Find similar Staff ML Engineer - AWS Trainium & SageMaker jobs: