Amazon

Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp

Amazon$182K — $247K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree
  • 8+ years in AI/ML, distributed computing, or GPU infrastructure
  • 3+ years in customer-facing ML architecture roles
  • 10+ years in IT development or consulting within software/cloud computing/AI
  • Experience with a major deep learning framework (PyTorch, TensorFlow, JAX)

Responsibilities

  • Lead technical deep-dives and performance optimizations for AI/ML workloads
  • Design and implement pioneering AI/ML training and inference solutions
  • Support customers in establishing critical HPC capabilities
  • Partner with service teams to boost model training throughput
  • Develop and nurture relationships with enterprise stakeholders as a trusted advisor

Benefits

  • Comprehensive health insurance including medical, dental, vision
  • 401(k) matching
  • Paid time off and parental leave
  • Mental health support and employee assistance programs
  • Adoption and surrogacy reimbursement
Full Job Description
Amazon Web Services (AWS) is seeking an experienced Principal AI/ML HPC Specialist to join our Technical Account Manager (TAM) team. You'll be at the forefront of solving complex AI HPC implementation challenges, guiding NAMER Resarch labs to enterprise customers through their most ambitious machine learning transformation journeys. By combining deep technical expertise with collaborative problem-solving, you'll help organizations unlock the full potential of artificial intelligence and machine learning technologies - from distributed model training on GPU clusters to production-grade inference at scale. AWS Support includes experts from across AWS who help our customers design, build, operate, and secure their cloud environments. Customers innovate with AWS Professional Services, upskill with AWS Training and Certification, optimize with AWS Support and Managed Services, and meet objectives with AWS Security Assurance Services. Our expertise and emerging technologies include AWS Partners, AWS Sovereign Cloud, AWS International Product, and AI/ML-native solutions. You'll join a diverse team of technical experts in dozens of countries who help customers achieve more with the AWS cloud. Key job responsibilities Deliver Strategic Technical Engagements - Lead comprehensive technical deep-dives and performance optimization for enterprise AI/ML workloads, including distributed training cluster architecture using AWS Parallel Computing Service (PCS) and AWS ParallelCluster, the latest GPU-accelerated computing (i.e. P6/P6e , G7/G7e instances), AWS Trainium-based training (Trn3 UltraServers), and multi-node NCCL communication tuning over EFA's SRD protocol. Architect and Validate Innovative Solutions - Design and implement production-grade AI/ML training and inference solutions leveraging Slurm-based job scheduling, distributed training frameworks (PyTorch FSDP, DDP, DeepSpeed, Megatron-LM), SageMaker HyperPod for managed GPU clusters with automated health checks and node replacement, high-performance parallel storage (Amazon FSx for Lustre), and container runtimes on Deep Learning AMIs (DLAMIs) against reference architectures and HPC lens to ensure performance, reliability, and cost governance at scale.. Architect solutions using P6e UltraServers for multi-trillion parameter frontier models and Trn3 with the AWS Neuron SDK for cost-optimized training and inference. Enable Customer Success - Support customers in implementing business-critical HPC capabilities, including the development of large language model (LLM) (Llama, GPT-class models), physics-informed neural networks (PINNs) and surrogate models, MLOps pipelines, simulation-ML hybrid architectures orchestrated by AWS Step Functions and AWS Batch, distributed data processing, cluster observability, and governance controls for GPU/Trainium-intensive workloads. Enable Business Critical Outcomes - Partner with with service teams to enhance model training throughput, optimize NCCL collective communications, improve GPU/Trainium utilization across multi-node UltraClusters, and drive operational efficiency through proactive monitoring, automated failure recovery (HyperPod health checks), and capacity planning (EC2 Capacity Blocks for ML). Contribute to product roadmap PFR, share refrerence architecture, performance , and benchmarks with broader TAM and Technical communities Serve as Trusted Advisor and Advocate - Develop and nurture technical partnerships with enterprise stakeholders, serving as the trusted advisor for AI/ML infrastructure decisions spanning compute, networking (Elastic Fabric Adapter with SRD), storage, orchestration, and the HPC-to-AI convergence journey. A day in the life Your day will be dynamic and impactful, involving deep technical consultations on distributed training architectures, strategic solution design for GPU and Trainium cluster deployments, and collaborative problem-solving across multi-node ML environments. You'll engage with technical leaders, architect innovative AI/ML implementations - from Slurm-managed PCS clusters and SageMaker HyperPod to PyTorch FSDP/DeepSpeed training jobs and Neuron SDK compilation workflows - and provide expert guidance that bridges machine learning infrastructure with business objectives. You will partner with TAMs, SAs, and service teams to provide customers with AWS AI/ML best practice guidance, diving deep into machine learning infrastructure services (PCS, ParallelCluster, HyperPod, Batch), promoting customers' AI/ML workloads to production, developing regional AI/ML strategies, advising on HPC-to-AI convergence patterns (simulation-surrogate loops, physics-informed neural networks), and training field teams on distributed training patterns, GPU/Trainium cluster operations, and the use cases and benefits of artificial intelligence and machine learning at scale. BASIC QUALIFICATIONS - Bachelor's degree - 8+ years of experience in AI/ML, distributed computing, or GPU-accelerated infrastructure (e.g., model training, inference systems, HPC for ML) - 3+ years of hands-on experience designing, implementing, or consulting on large-scale ML training or inference architectures in a customer-facing role - 10+ years of IT development or implementation/consulting in the software, cloud computing, or AI/ML industries - Experience with at least one major deep learning framework (PyTorch, TensorFlow, JAX) in a production or research environment - Demonstrated ability to serve as a trusted technical advisor to enterprise customers PREFERRED QUALIFICATIONS - Deep experience with distributed training techniques including data parallelism, model parallelism, pipeline parallelism, and Fully Sharded Data Parallel (PyTorch FSDP) - Experience with distributed training frameworks such as PyTorch DDP, DeepSpeed, and Megatron-LM for multi-node model training - Hands-on experience with GPU/accelerator cluster infrastructure: NVIDIA Blackwell (GB200, B200, B300), H100/H200 GPUs, AWS Trainium (Trn3/Trn2), NVLink/NVSwitch, InfiniBand or Elastic Fabric Adapter (EFA), and NCCL collective communications tuning - Experience with AWS Neuron SDK (torch-neuronx, neuronx-nemo-megatron) for compiling and optimizing models on Trainium and Inferentia (Inf2) instances - Familiarity with SageMaker HyperPod for managed distributed training clusters including automated health checks, node replacement, and checkpoint-based recovery - Experience with HPC job schedulers (Slurm, PBS, LSF) for orchestrating multi-node ML training workloads - Experience with high-performance parallel file systems (Amazon FSx for Lustre, GPFS/Spectrum Scale) for ML data pipelines - Familiarity with AWS Parallel Computing Service (PCS), AWS ParallelCluster, AWS Batch, or equivalent managed HPC/ML cluster services - Experience training or fine-tuning large language models (LLMs) such as Llama, GPT, or similar transformer architectures at multi-billion parameter scale - Understanding of HPC-AI convergence patterns: simulation-surrogate loops, physics-informed neural networks (PINNs), graph neural networks for molecular property prediction, and data format interoperability (HDF5, VTK, NetCDF to ML-ready tensors) - Knowledge of ML Ops tooling, container orchestration for training (Docker, Enroot, Pyxis), and Deep Learning AMIs (DLAMIs) - Experience with cluster observability and monitoring for GPU/Trainium utilization, training throughput, and job performance (CloudWatch, Prometheus, Grafana) - Experience with EC2 Capacity Blocks for ML, Capacity Reservations, or similar GPU capacity planning strategies - Experience with pipeline orchestration using AWS Step Functions for simulation-ML workflows - Experience with containers, EKS and ECS - Track record of driving operational excellence and proactive risk mitigation for mission-critical AI/ML workloads - AWS certifications (Solutions Architect Professional, Machine Learning Specialty) preferred The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits. USA, TX, Austin - 182,800.00 - 247,300.00 USD annually USA, TX, Dallas - 182,800.00 - 247,300.00 USD annually USA, VA, Herndon - 182,800.00 - 247,300.00 USD annually USA, WA, Seattle - 182,800.00 - 247,300.00 USD annually

About Amazon

Audible is a provider of spoken audio information and entertainment , on the Internet. They provide premium spoken audio content, such as audio versions of books and newspapers and radio programs, that is delivered over the Internet and played back on personal computers and hand-held electronic devices. The Audible service allows consumers to purchase and download their content from their Website, store it in digital files and play it back on personal computers and electronic devices. More than 15,000 hours of audio content are available on their Web site, including audio versions of books, periodicals and radio programs. Several manufacturers have agreed to support and promote the playback of their content on their hand-held audio-enabled electronic devices.

Amazon Careers

Joining Amazon presents an unparalleled opportunity to become part of a vibrant team pushing the boundaries of innovation and growth in the global marketplace. As a leader in e-commerce, technology, and logistics, Amazon offers a variety of job opportunities that cater to a range of skills and professional interests. Work You’ll Do At Amazon, every day is an opportunity to collaborate with the brightest minds in technology and business to redefine what’s possible. Whether you’re interested in software development, marketing, human resources, or customer service, Amazon has a position waiting for you. Transform the way the world shops and innovates with our diverse and inclusive team. Amazon is not just a company; it’s a community where you can drive real change and contribute to projects impacting millions globally. Lead with Innovation and Leadership Amazon is the perfect place to enhance your leadership and innovation skills. Our culture encourages pushing the envelope and imagining the unimaginable. Here, you will lead projects that challenge the status quo and define new industry standards. Work with a team that values diversity and is committed to creating an inclusive environment. Our leadership is focused on harnessing the collective power of unique perspectives to foster growth and innovation. Explore Amazon’s Employment Benefits Amazon’s commitment to its employees extends beyond just career growth. We offer competitive benefits, including health care, parental leave, and diversity training, ensuring that our team not only excels professionally but also enjoys well-being and security. Internship and Networking Opportunities Start your career with an Amazon internship and gain hands-on experience that matters. Our internships provide a gateway to full-time employment and an opportunity to network with professionals across various sectors of the company. Future-Proof Your Career With Amazon, your career path is filled with numerous opportunities for advancement. Our learning and development programs are designed to nurture your professional growth and keep you at the forefront of industry trends. Stay Connected Join Our Team Discover the job opportunities at Amazon that match your skills and interests. We are constantly on the lookout for passionate, curious, and innovative team players ready to make a difference. Keep Up to Date Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here. Job Alert Emails Customize your subscription to receive job alerts, the latest news, and insider tips tailored to your preferences. Explore the exciting and rewarding career opportunities that await at Amazon. Amazon is more than just a company—it’s a platform for building a promising future. Whether you’re starting or looking to advance your career, Amazon offers the resources, support, and network you need to succeed. Join us, and be a part of our continuing mission to be Earth's most customer-centric company.
Learn more about Amazon
Size
1,608 employees
Market Cap
$832.6 billion
Industry
Net Income
$21.3 billion
Founded
1994
5 Year Trend
+28.1%
Revenue
$386 billion
NASDAQ

Similar Jobs

More Jobs at Amazon

More Information Technology Jobs

Find similar Principal AI/ML HPC Specialist Technical Account Manager (STAM) , AWS Enterprise Support, NAMER-Sp jobs: