The Infrastructure Compute Site Reliability Engineering mission is to own and manage the successful operation of our underlying cell infrastructure system, along with elements of service discovery, secrets management and related software layers. We're looking for a skilled Senior Site Reliability Engineer with strong programming skills to help us build Roblox's private cloud, productionize our growing Kubernetes-based infrastructure, and institute reliability best practices across the Roblox Compute team.
You will:- Design and Develop systems & libraries that promote fault-tolerance and resilience, automate much of the management and lifecycle of our clusters, and ensure systems are observable.
- Promote and Institute reliability best practices across the Infra Compute group, drive common reliability initiatives. Provides collaborative technical reviews and operational guidance to strengthen system reliability.
- Build, Automate and Standardize process automation to create a "golden path" of tooling and platform support that powers the fundamental Roblox ecosystem.
- Create Tooling that provides production guardrails, by evaluating release candidate capacity with load testing tooling before deploying to production.
- Create Performance Monitoring Services and observability towards understanding capacity issues and platform degradations, monitoring production services and their changes, like generalized canarying services with alerting.
- Analyze systems and system designs for production readiness
You have:- A Bachelor degree (or equivalent professional experience) in Computer Science or related engineering field with a proven track record including at least 6 years as an SRE or Software Engineer.
- Fluency with high-level programming languages like Go, Java, C#.
- Experience with Kubernetes, or similar orchestration systems. Experience in Nomad, Vault, and Consul is strongly desired.
- Experience and good habits around building software and tools and getting them adopted. Your system's focus advises a view of code needing to be deeply reliable.
You are:- A Partner: You know that the best tools integrate broadly with the tooling ecosystem. You approach partners and processes with curiosity and seek to understand a problem deeply before you start coding.
- A Developer: You love building durable and reliable complex systems.
- Passionate about problem-solving, finding creative work solutions, and addressing unexpected challenges as part of a team.
- Problem Solver: You ask the right questions to tackle issues within your expertise and you use data to test your theories.
- Planner: You have experience in large project lifecycles. You have experience working in sprints, breaking down complex tasks into achievements, and reporting status to keep project scheduling accurate.
For roles that are based at our headquarters in San Mateo, CA: The starting base pay for this position is as shown below. The actual base pay is dependent upon a variety of job-related factors such as professional background, training, work experience, location, business needs and market demand. Therefore, in some circumstances, the actual salary could fall outside of this expected range. This pay range is subject to change and may be modified in the future. All full-time employees are also eligible for equity compensation and for benefits as described on
this page.
Annual Salary Range
$243,290-$295,250 USD
Roles that are based in an office are onsite Tuesday, Wednesday, and Thursday, with optional presence on Monday and Friday (unless otherwise noted).