Full Job Description
Netflix's Infrastructure Engineering organization provides the backbone that powers Netflix's global scale and operations across business verticals. Within this organization, Cloud Infrastructure Engineering provides a secure, scalable foundation that accelerates the productivity and innovation of Netflix
technology teams by abstracting infrastructure complexity behind a managed platform. Compute sits at the center of that foundation: every service, batch job, training run, inference request, and agent at Netflix ultimately runs on the compute platform we build.
Our Compute platform runs many kinds of workloads at a massive scale. That includes stateless and stateful services, batch jobs, streaming, media encoding, and a growing amount of distributed training and inference across our GPU fleet. Agentic workloads are now part of that mix. They introduce a new set of platform needs. Identity, sandboxed execution, observability, and cost accountability all need to be built in from the start.
The Opportunity
We’re looking for a Principal Product Manager with broad knowledge of distributed systems and infrastructure, strong systems thinking, and the ability to work with engineers at architectural depth. You’ll lead the modernization of Netflix’s Compute Platform as AI becomes a larger part of what we build. The goal is a common platform where services, batch jobs, training, inference, and agents can run securely, reliably, and efficiently at Netflix scale. You’ll own product direction and execution across the Compute portfolio, making the tradeoffs required to evolve our infrastructure, support new workload patterns, use capacity effectively, and meet specialized needs across Studio, Ads, and Games. You’ll turn deep technical context into a coherent strategy, sound product decisions, and meaningful results across the business. You’ll have the autonomy to make those decisions and the accountability that comes with them.
What You Will Do
• Set the strategic direction for Netflix’s Compute platforms in an agent-first world. You’ll define how our platform evolves to support this future, including guiding the move from Titus to a modern Kubernetes stack and weighing migration costs against the value of a common platform. At Netflix scale, this evolution is as much a product challenge as an engineering one.
• Improve how we manage GPU capacity. You’ll own the roadmap for provisioning, scheduling across regions, and fleet utilization. The goal is to give training and inference workloads the capacity they need within seconds while using the fleet efficiently.
• Define how agents run at Netflix. You’ll lead the development of a secure, managed environment that can support millions of agents, with built-in identity, isolation, governance, observability, and cost tracking.
• Make compute easier to use by creating a seamless experience across the full compute spectrum, from VMs to functions. You’ll improve our managed offerings so teams can focus on their work rather than infrastructure details, guided by regular engagement with service owners, ML engineers, and teams building agents.
• Improve efficiency across the fleet to drive hundreds of millions of dollars of impact. You’ll work on defining systems that manage compute supply and demand to ensure efficient resource use and that compute costs grow more slowly than demand. You’ll also establish clear measures for adoption, utilization, cost per workload, and reliability.
• Partner with teams across Ads, Live, Games, and Studio to ensure our compute platform supports their most demanding workloads, from sudden traffic spikes during ad breaks and live events to latency-sensitive processing in regions with limited local infrastructure.
What We're Looking For
Strong technical product management skills are table stakes. What will set someone apart in this role is:
• More than 10 years of experience in technical product management or a similar technical role in compute infrastructure. You should be comfortable debating architecture, trade-offs, and system quality with senior engineers.
• Deep knowledge of compute infrastructure, including containers, Kubernetes, scheduling and orchestration, autoscaling, capacity management, multi-tenancy, and cloud economics.
• A strong understanding of AI infrastructure. This includes GPU provisioning, scheduling, and utilization for distributed training and inference, as well as the needs of agentic workloads such as secure execution, workload identity, isolation, and cost attribution.
• Experience leading a large replatform or migration from start to finish. You understand that sequencing, adoption, incentives, and the long tail of migration are product decisions, not cleanup work for engineering.
• The ability to earn the trust of experienced engineers and challenge their thinking in useful ways.
• You make clear decisions with incomplete information, hold direction under pressure, and change course when the evidence calls for it.
• Strong leadership and communication. You can explain complex decisions to technical and executive audiences, build alignment across teams, and influence without relying on authority. You assume good intent, seek out different perspectives, give and receive direct feedback, and prioritize shared outcomes over individual ownership.
More About Compute @ Netflix
Generally, our compensation structure consists solely of an annual salary; we do not have bonuses. You choose each year how much of your compensation you want in salary versus stock options. To determine your personal top of market compensation, we rely on market indicators and consider your specific job family, background, skills, and experience to determine your compensation in the market range. The range for this role is $660,000.00 - $1,000,000.00. This compensation range will vary based on location.
Netflix provides comprehensive benefits including Health Plans, Mental Health support, a 401(k) Retirement Plan with employer match, Stock Option Program, Disability Programs, Health Savings and Flexible Spending Accounts, Family-forming benefits, and Life and Serious Injury Benefits. We also offer paid leave of absence programs. Full-time hourly employees accrue 35 days annually for paid time off to be used for vacation, holidays, and sick paid time off. Full-time salaried employees are immediately entitled to flexible time off. See more details about our Benefits here.
Netflix is a unique culture and environment. Learn more here.