The ML Capacity Delivery Team (MLZ) is seeking a Principal Technical Infrastructure Program Manager to define and drive the technical vision, strategy, and roadmap for the systems and tools that enable ML capacity delivery at global scale. In this role, you will serve as the single-threaded technical leader responsible for building, evolving, and scaling the platforms and automation that underpin how we plan, track, and deliver ML infrastructure; from demand signal through rack-level installation.
You will work at the intersection of infrastructure delivery, data engineering, and tooling while partnering with systems engineers, software development teams, Business Intelligence Engineers (BIE), operations leaders, and cross-functional stakeholders to translate complex operational requirements into scalable, reliable systems. The ideal candidate is a seasoned technical leader who thrives in ambiguity, can articulate a compelling 1-3 year technical roadmap, and has a track record of driving cross-organizational alignment to deliver platforms that fundamentally improve how teams operate at scale.
Key job responsibilities
• Own the technical vision and multi-year roadmap (1-3 years) for ML Capacity Delivery systems and tools, aligning investments to business priorities and capacity delivery goals across core and Gen AI platforms.
• Drive cross-functional alignment across systems engineering, software development, BIE, operations, and planning teams to define requirements, prioritize capabilities, and deliver integrated tooling solutions.
• Architect and deliver scalable platforms that automate and optimize capacity delivery workflows; including demand planning, build tracking, supply chain visibility, and operational reporting.
• Serve as the technical authority for systems and tools within MLZ, influencing Director/VP-level leadership on investment decisions, build-vs-buy trade-offs, and platform strategy through data-driven business cases.
• Identify and eliminate operational friction by mapping end-to-end workflows, quantifying inefficiencies, and delivering automation that reduces manual effort and accelerates delivery velocity.
• Establish engineering excellence; define standards for system reliability, data quality, observability, and documentation that enable teams to build and operate with confidence at scale.
• Force-multiply across the organization by creating reusable frameworks, mechanisms, and best practices that elevate the effectiveness of multiple teams and programs simultaneously.
• Drive clarity in highly ambiguous environments where the business strategy, architectural approach, and problem definition may not yet exist. While decomposing complex challenges into actionable execution plans.
• Build and maintain strategic partnerships with internal platform teams, data engineering organizations, and tooling providers to ensure MLZ systems integrate seamlessly with the broader AWS infrastructure ecosystem.
• Establish metrics, dashboards, and reporting mechanisms to provide stakeholders and senior leadership clear insight into platform health, delivery timelines, adoption, and business impact.
BASIC QUALIFICATIONS
- 10+ years of program management, tech program management, or related experience
- 10+ years of working directly with engineering teams experience
- Experience defining roadmap strategy and prioritizing deliverables for your team products
- 4+ years of designing or architecting (design patterns, reliability and scaling) of new and existing systems experience
- Bachelor's degree or above in Computer Science, Computer Engineering, Information Management, Information Systems, or other related discipline
- • Experience managing programs across cross-functional teams, building processes and coordinating release schedules
PREFERRED QUALIFICATIONS
- Experience working in data centers or critical infrastructure
- Experience in automation or monitoring frameworks, deployment or development
- Experience performing complex business case analysis to justify technical decisions and presenting the justification to management in a high-level review
- Experience communicating and presenting to senior leadership
- • Experience with AI/ML infrastructure, GPU/accelerator deployments, or high-performance computing environments
- • Track record of driving technical standardization and process improvement across globally distributed teams
- • MBA or Master's degree in Engineering, Computer Science, or related field
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, VA, Herndon - 176,900.00 - 239,400.00 USD annually
USA, WA, Seattle - 176,900.00 - 239,400.00 USD annually