Full Job Description
Ready to build the platforms that make AI at scale possible? NVIDIA is seeking a Senior Staff Platform Engineer to architect, build, and scale foundational infrastructure for some of our most demanding compute and AI/ML workloads. The role spans distributed systems, cloud, networking, content delivery, automation, and reliability engineering. Success means turning complex infrastructure challenges into resilient platforms that can grow with NVIDIA’s rapidly evolving needs. This is a hands-on technical leadership role with broad influence across Cloud, Networking, Security, AI/ML, and Developer Infrastructure teams.
What You Will Be Doing:
This role focuses on architecting, building, and scaling highly available platform services for AI/ML, distributed compute, and data-intensive workloads. The work spans cloud, compute, GPU infrastructure, networking, storage, DNS, load balancing, proxies, traffic management, and content delivery, with end-to-end ownership from architecture and design through implementation, production readiness, and large-scale adoption. A key part of the role is advancing CDN and edge infrastructure, including HTTP caching, origin architecture, TLS, WAF, rate limiting, traffic routing, and global load balancing. The role also drives automation, infrastructure-as-code, and self-service capabilities using technologies such as Python, Go, Kubernetes, and Terraform. Observability, capacity analytics, incident learnings, and performance data are used to continuously improve reliability, scalability, efficiency, and operational simplicity. Close collaboration with Cloud, Networking, Security, AI/ML, and infrastructure teams is central to solving complex problems that span multiple technology domains.
What We Need To See:
- Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, or equivalent experience
- 12+ years of relevant industry experience.
- Proven success architecting, building, and operating large-scale distributed platforms or infrastructure systems in production.
- Strong technical depth in several areas such as cloud infrastructure, distributed systems, networking, compute, storage, platform engineering, or content delivery.
- Deep knowledge of Linux/Unix, TCP/IP, DNS, TLS, HTTP/S, proxies, load balancing, availability, scalability, and fault-tolerant system design.
- Strong programming and automation skills with Python, Go, or similar languages, plus hands-on experience with infrastructure-as-code and orchestration.
- Experience with AWS, Azure, or Google Cloud Platform and the ability to troubleshoot complex systems across application, operating system, network, and infrastructure layers.
- Demonstrated ability to independently drive architecture and implementation across multiple teams, communicate effectively in complex situations, and mentor other engineers.
Ways To Stand Out From The Crowd:
- Experience building platforms for AI/ML training, inference, model serving, GPU-accelerated workloads, distributed compute, or high-performance computing.
- Deep expertise with CDN and edge platforms such as Akamai, AWS CloudFront, Fastly, or Cloudflare, including caching, origin design, WAF, DNS, TLS, and global traffic management.
- Experience developing self-service platform capabilities that enable engineering teams to consume infrastructure reliably and at scale.
- Proven use of SLIs, SLOs, error budgets, capacity analytics, and reliability metrics to deliver measurable improvements.
- Experience distributing models, datasets, containers, software artifacts, or other large objects across globally distributed environments.
Want to work on infrastructure where scale, AI, networking, and reliability come together? Join us!
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 200,000 USD - 322,000 USD.
You will also be eligible for equity and .
Applications for this job will be accepted at least until August 23, 2026.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.