Riot Games

Senior Software Engineer, Platform & Infrastructure - Riot Technology

Riot Games$135K — $160K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science or related field, or equivalent experience.
  • 3+ years of software engineering experience in infrastructure, platform engineering, or SRE roles.
  • Experience operating distributed systems in production environments.
  • Strong experience with Kubernetes, AWS or GCP, and infrastructure-as-code.
  • Experience with GPU compute infrastructure and resource optimization for training workloads.
  • Proficiency in Python and understanding of networking and distributed systems.

Responsibilities

  • Build and operate Kubernetes and multi-node GPU clusters for ML training.
  • Design infrastructure for scalable simulation environments for data collection and evaluation.
  • Build CI/CD and deployment automation across cloud environments.
  • Improve platform reliability, cost efficiency, and operational maturity.
  • Create observability and monitoring systems for ML workloads.
  • Develop internal APIs and developer tooling for training workflows.
  • Support MLOps workflows including automated training and model management.

Benefits

  • Open paid time off policy promoting work/life balance.
  • Flexible work schedules.
  • Medical, dental, and life insurance coverage.
  • Parental leave for employees and their families.
  • 401k with company match.
Full Job Description
Platform and infrastructure engineers at Riot build the foundational systems that enable teams to develop, deploy, and operate systems at global scale. They partner across disciplines with other engineers, data scientists, designers, and product teams to ensure reliable, scalable, and secure infrastructure underpins every ML capability that reaches players.

As a Senior Platform & Infrastructure Engineer on the Riot Technology team, you will design, build, and operate the core infrastructure and ML platforms. Your focus will be on the computing and orchestration platforms that power large-scale distributed training of agents (e.g. via RL, IL, and other techniques), simulation environments, and policy evaluation, as well as the CI/CD, infrastructure-as-code, observability, and developer tooling that keep these systems production-grade. You will close critical infrastructure gaps across the team's stack, driving improvements to standards, automation, and operational maturity. You will operate independently on multi-month work efforts and begin to influence technical direction beyond your immediate team. You will report to the Manager of Machine Learning.

Responsibilities:
  • Build and operate Kubernetes, multi-node GPU clusters, and networking infrastructure for distributed ML bot training and large-scale policy evaluation.
  • Design infrastructure for running simulation environments at scale, enabling parallel rollouts, data collection, training, and evaluation.
  • Build CI/CD, deployment automation, artifact management, and infrastructure-as-code across cloud environments.
  • Improve platform reliability, cost efficiency, performance, reproducibility, auditability, and operational maturity.
  • Build observability, monitoring, alerting, health indicators, and SLO-aligned dashboards for infrastructure and ML workloads.
  • Develop internal APIs, control planes, templates, and developer tooling for distributed training and evaluation workflows.
  • Support MLOps workflows including automated training pipelines, model artifact management, experiment tracking, and reproducible ML lifecycle operations.
  • Build security and governance controls, manage production incidents, drive root-cause remediation, mentor engineers, and support recruiting for platform roles.

Required Qualifications:
  • Bachelor's degree in Computer Science or a related field, or equivalent practical experience.
  • 3+ years of software engineering experience, with meaningful experience in infrastructure, platform engineering, or SRE roles.
  • Experience operating distributed systems in production and keeping them healthy under real load.
  • Strong experience with Kubernetes, AWS or GCP, infrastructure-as-code, CI/CD, deployment automation, and production tooling.
  • Experience with GPU compute infrastructure, including scheduling, multi-node orchestration, and resource optimization for long-running training workloads.
  • Proficiency in Python and solid understanding of networking, microservices, core infrastructure services, and distributed systems fundamentals.

Desired Qualifications:
  • Familiarity with MLOps workflows such as model versioning, pipeline orchestration, experiment tracking, artifact management, and reproducible ML workflows.
  • Experience with distributed training or HPC frameworks, inference serving, systems languages, high-performance networking, Unreal/client-server architecture, AI-assisted development tools
  • Passion for games and player experience.

For this role, you'll find success through craft expertise, a collaborative spirit, and decision-making that prioritizes the delight of players. We will be looking at your past studies, experience, and your personal relationship with games. If you embody player empathy and care about players' experiences, this could be your role!
Our Perks:

Riot focuses on work/life balance, shown by our open paid time off policy and other perks such as flexible work schedules. We offer medical, dental, and life insurance, parental leave for you, your spouse/domestic partner, and children, and a 401k with company match. Check out our benefits pages for more information.

At Riot Games, we put players first. That mission drives every decision in our quest to create games and experiences that make it better to be a player. Whether you're working directly on a new player-facing experience or you're supporting the company as a whole, everyone at Riot is part of our mission. And just like in our games, we're better when we work together. Our goal is to create collaborative teams where you are empowered to bring your unique perspective everyday. If that sounds like the kind of place you want to work, we're looking forward to your application.

About Riot Games

Riot Games is a video game developer and publisher based in Los Angeles, California. The company was founded in 2006 by Brandon Beck and Marc Merrill, and is best known for its flagship game, League of Legends. The game has become one of the most popular esports titles in the world, with millions of players and fans around the globe. Riot Games is committed to creating high-quality, immersive gaming experiences that bring people together and foster a sense of community. The company is also dedicated to promoting diversity and inclusion in the gaming industry, and has launched several initiatives to support underrepresented groups.
Learn more about Riot Games
Size
2,500 employees
Industry
Founded
2006

Similar Jobs

More Jobs at Riot Games

More Information Technology Jobs

Find similar Senior Software Engineer, Platform & Infrastructure - Riot Technology jobs: