Amazon

Sr Technical Program Manager - Hardware, AWS Generative AI & ML Servers

Amazon • $148K — $201K *
Telecommunications & Hardware
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering, or related discipline.
  • 6+ years of technical product or program management experience.
  • 5+ years of experience driving hardware development programs through the product lifecycle.
  • Experience managing programs across cross-functional teams and coordinating release schedules.
  • Preferred Master's degree in relevant fields and experience with GPU server platforms.

Responsibilities

  • Define and influence program strategy, objectives, and success criteria aligned with organizational goals.
  • Build mechanisms for program visibility, including metrics and dashboards for early risk detection.
  • Streamline delivery processes, eliminating dependencies and redundant coordination overhead.
  • Facilitate requirements gathering and develop Technical Requirements Documents (TRDs).
  • Manage cross-functional alignment through design reviews and readiness assessments for deployment.
  • Track ODM partnerships, ensuring quality checkpoints and readiness across builds.
  • Drive root cause analysis for hardware failures and define acceptance criteria for qualification testing.

Benefits

  • Comprehensive health insurance (medical, dental, vision).
  • 401(k) matching program.
  • Paid time off and parental leave.
  • Support for mental health, adoption, and surrogacy reimbursement.
  • Flexible spending accounts (FSA).
Full Job Description
AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML workloads at cloud scale. Our team designs, builds, and operates this fleet - solving systemic hardware issues and building systems that detect and prevent recurrence so customers experience the highest quality of service.

We are seeking a Senior Technical Program Manager to drive end-to-end delivery of GPU-accelerated servers across our global fleet. You will coordinate cross-functional engineering teams spanning hardware, firmware, and software, manage ODM partnerships across multiple continents, and establish closed-loop quality systems that drive continuous improvements. This role requires technical depth to translate engineering constraints into program risk, combined with program management excellence to deliver complex hardware at global scale.

What You Will Do

You will own programs where the critical path runs through silicon, firmware, and software teams simultaneously. You will translate ambiguity into structure: turning a fleet telemetry signal into a corrective action plan with quantified failure rates, a customer requirement into a new platform milestone with EVT/DVT/PVT gates, or a manufacturing escape into a design change with updated validation criteria. You will drive decisions on program trade-offs - adjusting scope when qualification gates slip, balancing deployment speed against fleet risk, and determining when to accept interim mitigations instead of holding for root-cause fixes. When a large scale of GPU servers depend on your program landing on time, you are the one ensuring hardware readiness, qualification completeness, and operational handoff happen without gaps.

The Ideal Candidate

You have deep technical intuition across hardware and software - enough to challenge engineering decisions, not just track them. You thrive in ambiguity, bringing structure to programs where requirements, timelines, and dependencies are still forming. You align priorities across teams in different organizations, and you escalate with data, not noise. You actively mentor and develop others - TPMs and engineers alike. You contribute to hiring, promotion assessments, and raising the bar for program management practices in your organization.

Key job responsibilities

Strategy & Mechanisms

* Define program strategy, objectives, and success criteria; influence resource allocation and priority decisions across engineering workstreams to align with organizational goals

* Build and own mechanisms for program visibility - defining metrics, dashboards, and review cadences that enable data-driven decisions and early risk detection

* Streamline delivery processes across teams; identify and eliminate dependencies, redundant gates, or coordination overhead that slow velocity

Requirements & Planning

* Facilitate requirements gathering with internal customers; develop Technical Requirements Documents (TRDs) covering server specs, rack configurations, PCIe topology, power/cooling topology, and SKU definitions

* Build program timelines aligned to different phases (Program Initiation, Design, Qualification, Pilot, Post-launch) with critical path analysis, risk identification with new & unique changes to the hardware, and milestone tracking for the program.

* Drive Program Initiation reviews: scope definition, preliminary annualized failure rate predictions, resource planning, supply chain long-lead identification, and RFP (Request for Proposal) issuance to ODMs

Execution & Coordination

* Drive cross-functional alignment across hardware, firmware (BIOS, BMC, CPLD), software, and operations teams through design reviews, manufacturing readiness and production readiness for fleet deployment.

* Manage ODM partnerships: track EVT/DVT/PVT builds, manufacturing readiness gates, Bill of Materials (BOM) management in PLM systems, and quality checkpoints

* Identify blockers early, escalate dependencies before they impact critical path, and facilitate technical trade-off decisions across engineering workstreams.

* Communicate program status to leadership with clear reporting on milestone progress, risk posture, and mitigation plans.

Risk & Quality

* Challenge technical workstreams to surface risks early; quantify impact to schedule and reliability (annualized failure rate targets, availability SLAs) with proposed mitigations

* Drive root cause analysis of fleet-wide hardware failures and ensure corrective actions flow back into qualification criteria and design requirements

* Define acceptance criteria, coordinate qualification testing at server and rack levels, and manage go/no-go decisions for mission-critical AI/ML workloads

Transition

* Conduct knowledge transfer and document lessons learned; ensure operational readiness including automation, monitoring, and runbook completeness for production handoff

May require occasional (

A day in the life

You start the day syncing with ODM partners across time zones on build status and open engineering actions. Mid-morning, you run an engineering review connecting firmware, software, and hardware teams to unblock a qualification gate. In the afternoon, you triage a fleet reliability issue with operations data, drive alignment on corrective actions, and update executive stakeholders on program risk posture. You end the day reviewing NPI milestone readiness and ensuring the next design review has clear entry criteria.

About the team

The Hardware Engineering AI/ML UltraServer platform team is a group of engineers and technical program managers directly responsible for launching GPU-accelerated servers into the AWS fleet. Located in Seattle, Austin, and Cupertino, we collaborate with global development teams and ODM partners to deliver next-generation AI/ML infrastructure deployed in datacenters worldwide. We move fast with small, empowered teams delivering end-to-end - from server conception through fleet-scale operations.

BASIC QUALIFICATIONS

- Bachelor's degree in Computer Science, Electrical Engineering, Computer Engineering or a related discipline or equivalent

- Experience managing programs across cross-functional teams, building processes and coordinating release schedules

- 6+ years of technical product or program management experience, working directly with multiple engineering teams.

- 5+ years of experience driving hardware development programs (servers, racks, networking, or storage) through full product lifecycle including design, validation, manufacturing, and fleet deployment

PREFERRED QUALIFICATIONS

- Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or related fields

- Experience leading the design, automation, deployment, and support of large-scale infrastructure

- Experience facilitating discussions with senior leadership regarding technical / architectural trade-offs, best practices, and risk mitigation

- 5+ years of experience coordinating complex server programs with ODM/JDM partners across multiple geographies, managing design reviews, manufacturing readiness gates, and quality checkpoints

- 5+ years of experience working with GPU/accelerator server platforms, NPI (New Product Introduction), or hardware qualification programs across various phases (EVT, DVT, PVT)

- Experience driving root cause analysis and corrective action processes for hardware reliability issues at fleet scale

- Knowledge of server hardware development lifecycle: electrical/mechanical design, firmware, thermal/power validation, and manufacturing testing.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, CA, Cupertino - 171,000.00 - 231,400.00 USD annually

USA, TX, Austin - 148,700.00 - 201,200.00 USD annually

USA, WA, Seattle - 148,700.00 - 201,200.00 USD annually

About Amazon

Audible is a provider of spoken audio information and entertainment , on the Internet. They provide premium spoken audio content, such as audio versions of books and newspapers and radio programs, that is delivered over the Internet and played back on personal computers and hand-held electronic devices. The Audible service allows consumers to purchase and download their content from their Website, store it in digital files and play it back on personal computers and electronic devices. More than 15,000 hours of audio content are available on their Web site, including audio versions of books, periodicals and radio programs. Several manufacturers have agreed to support and promote the playback of their content on their hand-held audio-enabled electronic devices.

Amazon Careers

Joining Amazon presents an unparalleled opportunity to become part of a vibrant team pushing the boundaries of innovation and growth in the global marketplace. As a leader in e-commerce, technology, and logistics, Amazon offers a variety of job opportunities that cater to a range of skills and professional interests. Work You’ll Do At Amazon, every day is an opportunity to collaborate with the brightest minds in technology and business to redefine what’s possible. Whether you’re interested in software development, marketing, human resources, or customer service, Amazon has a position waiting for you. Transform the way the world shops and innovates with our diverse and inclusive team. Amazon is not just a company; it’s a community where you can drive real change and contribute to projects impacting millions globally. Lead with Innovation and Leadership Amazon is the perfect place to enhance your leadership and innovation skills. Our culture encourages pushing the envelope and imagining the unimaginable. Here, you will lead projects that challenge the status quo and define new industry standards. Work with a team that values diversity and is committed to creating an inclusive environment. Our leadership is focused on harnessing the collective power of unique perspectives to foster growth and innovation. Explore Amazon’s Employment Benefits Amazon’s commitment to its employees extends beyond just career growth. We offer competitive benefits, including health care, parental leave, and diversity training, ensuring that our team not only excels professionally but also enjoys well-being and security. Internship and Networking Opportunities Start your career with an Amazon internship and gain hands-on experience that matters. Our internships provide a gateway to full-time employment and an opportunity to network with professionals across various sectors of the company. Future-Proof Your Career With Amazon, your career path is filled with numerous opportunities for advancement. Our learning and development programs are designed to nurture your professional growth and keep you at the forefront of industry trends. Stay Connected Join Our Team Discover the job opportunities at Amazon that match your skills and interests. We are constantly on the lookout for passionate, curious, and innovative team players ready to make a difference. Keep Up to Date Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here. Job Alert Emails Customize your subscription to receive job alerts, the latest news, and insider tips tailored to your preferences. Explore the exciting and rewarding career opportunities that await at Amazon. Amazon is more than just a company—it’s a platform for building a promising future. Whether you’re starting or looking to advance your career, Amazon offers the resources, support, and network you need to succeed. Join us, and be a part of our continuing mission to be Earth's most customer-centric company.
Learn more about Amazon
Size
1,608 employees
Market Cap
$832.6 billion
Industry
Net Income
$21.3 billion
Founded
1994
5 Year Trend
+28.1%
Revenue
$386 billion
NASDAQ

Similar Jobs

More Jobs at Amazon

More Telecommunications & Hardware Jobs

Find similar Sr Technical Program Manager - Hardware, AWS Generative AI & ML Servers jobs: