Job descriptionAs a Principal Technical Program Manager, Bare Metal GPU & AI Cloud Delivery, you will own the end-to-end delivery of strategic AI Cloud programs, from infrastructure planning, capacity readiness, and bare metal GPU cluster bring-up through production launch, customer onboarding, and operational handoff. You will bring structure, execution discipline, and technical depth to programs spanning GPU systems, high-performance networking, storage, data center readiness, platform software, security, and customer-facing delivery.
You will partner closely with infrastructure engineering, AI cloud software engineering, data center operations, networking, security, supply chain, finance, solution architecture, product, and customer-facing teams to deliver reliable, scalable, production-ready GPU cloud infrastructure for AI training and inference workloads at scale.
This role requires a senior owner who can operate across hardware, data center, and cloud software domains; drive clarity in ambiguous environments; anticipate risks across vendors, facilities, and engineering teams; and ensure that every stage of AI Cloud delivery is planned, tracked, validated, and communicated with executive-level precision.
Job requirements• Bachelor's degree in a technical discipline or equivalent experience.
• 10+ years of experience in technical program management, cloud infrastructure, data centers, hardware infrastructure, or software engineering programs.
• Proven experience owning end-to-end delivery of complex infrastructure programs from planning and requirements through deployment, production readiness, customer launch, and operational handoff.
• Strong understanding of bare metal GPU infrastructure, GPU server platforms, high-performance networking, storage, cloud platforms, distributed systems, and data center dependencies such as power, cooling, rack layout, and network readiness.
• Experience leading GPU cluster bring-up, capacity enablement, hardware deployment, network integration, validation, and production launch across internal teams, vendors, and external partners.
• Ability to manage critical paths, risks, dependencies, schedules, decision forums, executive communications, and customer-facing delivery commitments in high-ambiguity environments.
Preferred Qualifications
• Experience in hyperscale cloud, AI infrastructure, GPU cloud, bare metal cloud, or HPC environments.
• Experience supporting NVIDIA-based AI platforms, high-density GPU clusters, InfiniBand or high-performance Ethernet fabrics, and cloud capacity bring-up with external partners.
• Familiarity with Kubernetes, cloud-native platforms, provisioning, orchestration, automation, observability, fleet management, and infrastructure-as-code practices.
• Experience partnering with solution architects, customer engineering, sales engineering, managed service providers, system integrators, and data center delivery teams.
• PMP, PgMP, Agile, Scrum, or similar certifications.
Job responsibilities• Own and drive end-to-end delivery processes for Bare Metal GPU and AI Cloud programs, from requirements definition, infrastructure design, capacity planning, procurement readiness, and deployment planning through commissioning, software deployment, customer handover, and operational steady state.
• Define, standardize, and scale repeatable delivery processes across multiple data centers, ensuring each site follows clear playbooks, milestones, dependency tracking, readiness criteria, risk management, escalation paths, and handoff procedures.
• Lead cross-functional execution across infrastructure design, data center operations, supply chain, logistics, infrastructure installation, cabling, networking, commissioning, cloud software deployment, security, customer engineering, and support teams.
• Partner with infrastructure design teams to translate AI Cloud capacity, GPU cluster architecture, power, cooling, rack layout, network fabric, storage, and operational requirements into executable multi-site delivery plans.
• Collaborate with supply chain and vendor partners to align GPU systems, network equipment, racks, optics, storage, firmware, spares, logistics, and site delivery schedules with program milestones and customer commitments.
• Coordinate infrastructure installation and site readiness activities, including rack placement, power and cooling validation, network turn-up, cabling completion, hardware acceptance, and issue remediation across internal teams and external contractors.
• Drive commissioning and production readiness reviews for each GPU cluster, ensuring hardware, firmware, networking, storage, automation, observability, security controls, support processes, and service acceptance criteria are fully validated before launch.
• Lead software deployment readiness across provisioning, orchestration, monitoring, capacity management, customer onboarding workflows, and platform service enablement to ensure AI Cloud environments are production-ready.
• Own customer handover planning and execution, including launch readiness, acceptance criteria, documentation, support transition, known-issue tracking, stakeholder communications, and post-launch stabilization.
• Establish portfolio-level governance for concurrent data center and AI Cloud delivery programs, providing executive-level reporting on progress, risks, dependencies, escalations, site readiness, customer impact, and business outcomes.
• Identify bottlenecks across process, tooling, vendor execution, installation workflows, commissioning, software deployment, and operational handoffs; drive durable improvements that increase deployment velocity, repeatability, quality, reliability, and cost efficiency.
Job benefits At IREN, we offer a highly competitive compensation package that includes base salary, annual performance incentives, and opportunities to build long-term wealth through equity programs. These offerings are part of our broader Total Rewards package, thoughtfully designed to support your health, well-being, and long-term success.
Compensation
- Salary range (for San Francisco, California location): USD $200,000 - 260,000/annum, depending on experience
- Short-term and Long-term Incentive Programs
Health & Wellness
- Medical, dental, and vision insurance coverage - 100% company paid for employees, 75% company paid coverage for dependents
- Company-paid life and disability insurance
- Voluntary life, critical illness, and accident coverage available
- Health Savings Accounts (HSA) - when combined with the High-Deductible Health Plan
- Employee Assistance Program and wellness resources
Financial Well-Being
- 401(k) retirement plan with company match
- Financial wellness tools and resources
Time Off & Flexibility
- Paid Time Off (PTO) and paid holidays
- Flexible work arrangements
Growth & Development
- Opportunities for advancement and internal mobility
- Training and personal development opportunities
Lifestyle & Culture
- Company events and team-building activities
We value diverse perspectives and believe that skills can be developed. If you're passionate about this role, we want to hear from you - whether you meet every criteria or not. Your unique experiences might be exactly what we need!