Wonder

Staff Site Reliability Engineer

Wonder • $208K — $216K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in SRE, DevOps, or infrastructure engineering with demonstrated Staff-level scope.
  • Deep experience with Infrastructure as Code - specifically Terraform.
  • Kubernetes expertise at the operator level, including EKS lifecycle management.
  • Knowledge of multi-region AWS architecture, with failover design capabilities.
  • Proficient with CI/CD tools such as Jenkins or GitHub Actions.
  • Experience with programming in Python, Go, or similar languages.
  • Familiarity with datastores like MySQL, MongoDB, and message brokers.

Responsibilities

  • Architect resilient, self-healing systems ensuring platform scalability.
  • Manage AWS infrastructure as code, handling Terraform and related tools.
  • Oversee Kubernetes platform lifecycle, including upgrades and autoscaling.
  • Drive observability improvements closing gaps before they lead to incidents.
  • Design strategies for handling seasonal traffic, especially during back-to-school.
  • Maintain CI/CD pipelines and essential deployment tools.
  • Lead incident management processes, including response and postmortems.

Benefits

  • Competitive compensation package with equity options.
  • 401(k) plan participation.
  • Choice of medical, dental, and vision plans.
  • Company paid short and long term disability coverage.
  • Flexible paid time off for exempt employees and paid vacation for non-exempt employees.
  • Paid sick leave in compliance with applicable law.
  • Paid parental leave and employee discounts on meals.
Full Job Description
About The Opportunity

This role is crucial for simplifying the dining experience for students across the US. You will be instrumental in architecting resilient and self-healing solutions, managing AWS infrastructure, closing observability gaps, designing scaling approaches, and shaping incident management processes. Your contributions will span the entire development lifecycle, encompassing the building and maintenance of CI/CD pipelines. Your role will be pivotal in ensuring the platform's scalability to support Grubhub's continuously expanding customer base, evidenced by the addition of 30 new campuses and a 25% year-over-year increase in order volume.

The Impact You'll Make
  • Architect resilient, self-healing systems and co-own the design of critical production services.
  • Own multi-region resilience: active-standby architecture, regional failover readiness, runbooks and drills, RTO/RPO targets, and the data-layer replication behind them (ElastiCache Global Datastore, MongoDB/Atlas, RDS).
  • Own AWS infrastructure as code - Terraform/Terraspace, Helm and Helmfile - from design through rollout.
  • Own the Kubernetes platform: EKS lifecycle, controller and add-on upgrades, ingress/gateway, and autoscaling (HPA/KEDA).
  • Own the observability platform end to end - logging, metrics, tracing and alerting pipelines - including its signal quality and its cost.
  • Drive reliability improvements using SLOs and telemetry data, closing observability gaps before they turn into incidents.
  • Design scaling and capacity strategy against a strongly seasonal traffic profile that peaks at back-to-school.
  • Own cloud cost accountability for the platform: right-sizing, reservations, and tracking realized savings.
  • Build and maintain CI/CD pipelines and the deployment tooling the team depends on.
  • Operate within PCI-scoped environments, respecting segregated clusters and access paths.
  • Shape incident management: lead incident response, postmortems, failure analysis, service reviews and architecture review.


What You'll Bring To The Table

Experience:
  • Demonstrated Staff-level scope: has owned a production platform end to end, driven technical direction across teams without formal authority, and led both incident response and architecture review. Typically 8+ years in SRE, DevOps or infrastructure engineering, though scope and impact weigh more than tenure.

Technical Skills:
  • Deep experience with Infrastructure as Code - Terraform (with Terraspace or a similar wrapper) - owning modules, state and rollout across multiple environments.
  • Kubernetes at operator depth: EKS lifecycle, Helm and Helmfile, controllers and add-ons, ingress/gateway, and autoscaling (HPA/KEDA).
  • Multi-region AWS architecture, including failover design and the data-layer replication it depends on.
  • Deep knowledge of CI/CD tools (e.g., Jenkins, GitHub Actions).
  • Software engineering experience in Python, Go, or a similar object-oriented language.
  • Proficiency with datastores (MySQL, MongoDB/Atlas, Redis/ElastiCache) and message brokers (RabbitMQ/AmazonMQ, SQS).
  • Experience with Microservice Architecture and Application Design.
  • Distributed monitoring experience, including SLOs, metrics, tracing and log pipelines.
  • Strong working knowledge of cloud fundamentals (AWS compute/containers, storage, Linux, networking).
  • Comfort operating in compliance-scoped environments (PCI or equivalent).

Soft Skills:
  • Strong technical writing, documentation, and communication skills.
  • Experience with highly trafficked web-based services.
  • Ability to set technical direction and build consensus across engineering teams in multiple time zones.


New York Base Salary: $208,500-$216,500.

Wonder uses geographic-specific salary structures, which means the salary offered may vary depending on where the job is located. The final salary offer will take into account various factors, such as the candidate's skills, education, training, credentials, and experience.

Benefits

The benefits applicable to this role include a competitive compensation package with equity and a 401(k). We also offer a choice of medical, dental, and vision plans, company paid short and long term disability coverage, paid time off including flexible time off for exempt employees, paid vacation for non-exempt employees, and paid sick leave in compliance with applicable law in addition to paid parental leave, discounted meals and exclusive perks across the Wonder family of brands.

Eligibility, effective dates, and available plan options vary by employment classification and location. To learn more about benefits for this role, visit our Careers page here.

About Wonder

World of Wonder Productions is an American production company founded in 1991 by filmmakers Randy Barbato and Portsmouth-born Fenton Bailey. Based in Los Angeles, California, the company specializes in documentary television and film productions, with credits including the Million Dollar Listing docuseries, RuPaul's Drag Race, and the documentary films Mapplethorpe: Look at the Pictures and The Eyes of Tammy Faye. Together, Bailey and Barbato have produced programming through World of Wonder for HBO, Bravo, HGTV, Showtime, the BBC, Netflix, and VH1. World of Wonder is perhaps best known for its contributions towards LGBTQ programming, for which they won an Outfest Annual Achievement Award in 2011. Their most well known LGBTQ production is RuPaul's Drag Race, having managed the career of drag queen and titular host RuPaul for many years before this, eventually producing the franchise alongside the majority of its live shows, podcasts, television specials, and conventions.
Learn more about Wonder
Industry
Founded
2015

Similar Jobs

More Jobs at Wonder

More Information Technology Jobs

Find similar Staff Site Reliability Engineer jobs: