Staff Site Reliability Engineer-Production Operations

Rivian and Volkswagen Group Technologies

$150K — $180K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years experience in SRE, systems engineering, or large-scale distributed platforms
  • Proven ability to coordinate high-severity incidents and communicate with executives
  • Strong foundation in distributed systems, networking, and cloud infrastructure
  • Hands-on coding skills in Python, Go, or similar for tooling and automation
  • Expertise in observability practices including metrics, logging, and SLO design
  • Familiarity with modern reliability practices and learning frameworks like Blameless reviews
  • Experience mentoring engineers and an interest in people management

Responsibilities

  • Coordinate incident response and streamline communication during crises
  • Facilitate post-incident reviews based on Learning From Incidents principles
  • Manage and prioritize action items from incident reviews with TPMs and dev teams
  • Measure and verify the effectiveness of implemented fixes
  • Design and implement automation and observability systems
  • Set technical standards and support the growth of the ProdOps team

Benefits

  • Base salary with eligibility for annual performance bonuses
  • Opportunity for equity participation
  • Flexible benefits tailored to the local market
  • Access to global benefits across the company
  • Potential for professional development and training opportunities
Full Job Description
Role Summary

We are looking for an SRE Lead to serve as the senior technical leader and player-coach for ProdOps. This is a hybrid role: you will set the technical direction of the team and lead from the front during incidents, while also growing and managing a small group of exceptional engineers as the function scales.

As the calm center during a crisis, you will maintain a high-level mental model of the entire production ecosystem, freeing engineers to focus strictly on debugging and mitigation. You will own the reliability feedback loop end to end, which includes running blameless post-incident reviews, coordinating major incidents across Cloud systems, vehicle development pipelines, and Product Security, driving the systemic action-item backlog with TPMs and development teams, and verifying that fixes hold in production.

This role suits a senior systems engineer who still writes code, thinks in terms of systems and failure modes rather than single root causes, and wants to build a lean, automation-first reliability practice rather than staff a support queue.

Responsibilities
Lead incident coordination and communication

Drive incident progress and coordinate cross-functional response as the central nervous system during a crisis. Own executive, customer, and parent-company communications, providing production expertise and communication leadership so engineers can concentrate on the technical problem. Maintain an accurate, high-level model of the full production ecosystem spanning Cloud, vehicle development pipelines, and Product Security.
Facilitate modern, blameless post-incident learning

Run post-incident reviews using Learning From Incidents (LFI) principles and HOWIE-style reporting. Move the organization away from the search for a single root cause and toward understanding how tooling, context, and multiple latent conditions combined to produce failure. Surface weaknesses in observability, process, testing, and tooling that impaired our ability to detect, mitigate, and recover.
Own and prioritize systemic action items

Manage the backlog of action items generated by reviews. While ProdOps does not write the fixes itself, you will prioritize, track, and drive these items to closure in partnership with TPMs and development teams, keeping leadership focused on customer impact.
Measure and verify efficacy

Once the fixes ship, measure and verify that they actually prevent recurrence. Drive accountability for outcomes and feed the results back into the Novel Incident Rate.
Build the automation and observability backbone

Design and build the systems, tooling, and AI-agent workflows that automate incident triage and administrative toil. Improve observability, including instrumentation, alerting, dashboards, and SLOs, so incidents are detected faster and understood more deeply. Write and review code where it multiplies the team's impact.
Lead and grow the team

Set technical standards and operating rhythm for ProdOps. Mentor and develop engineers, and as the team scales, take on hiring and people management while preserving the minimal-headcount, maximum-automation philosophy.

Qualifications

Minimum Qualifications:
  • Substantial experience in SRE, production/platform engineering, or systems engineering for large-scale distributed systems, including senior technical leadership or lead responsibilities.
  • Proven incident command experience: you have coordinated high-severity, cross-functional incidents and led communication with executives and external stakeholders under pressure.
  • Strong systems engineering fundamentals, including distributed systems, networking, cloud infrastructure, and an instinct for how complex systems fail.
  • Hands-on coding ability (e.g., Python, Go, or similar) sufficient to build automation, tooling, and integrations. This is not a code-free management role.
  • Deep observability expertise: instrumentation, metrics, logging, tracing, alerting, dashboards, and SLO/SLI design (Datadog or comparable platforms).
  • Fluency in modern reliability and post-incident practice: blameless reviews, Learning From Incidents (LFI), HOWIE, and systemic (non-single-root-cause) analysis.
  • A demonstrated bias toward eliminating toil through automation, and enthusiasm for using AI agents as force multipliers.
  • Excellent written and verbal communication; able to translate technical detail for executive and cross-company audiences.
  • Experience mentoring engineers, with the judgment and appetite to grow into formal people management.

Preferred Qualifications
  • Experience across both cloud services and hardware/vehicle or embedded development pipelines.
  • Exposure to Product Security or working closely with security teams during incidents.
  • Track record building an SRE or reliability function from an early stage.
  • Experience deploying LLM- or agent-based automation into production operational workflows.
Total Rewards

Full-time positions include base salary, eligibility for an annual performance bonus, and eligibility for equity.

In addition to base salary, Rivian and Volkswagen Group Technologies offers benefits tailored to the local market. For more information on the benefits available for full-time employees, check out our Global Benefits Site.

External candidates can apply for this role through the Rivian and Volkswagen Group Technologies careers site (https://rivianvw.tech/#careers). If you are a current employee, please apply through our internal job board.

Similar Jobs

More Jobs at Rivian and Volkswagen Group Technologies

More Information Technology Jobs

Find similar Staff Site Reliability Engineer-Production Operations jobs: