Senior Site Reliability Engineer (SRE)

UJET

• $140K — $180K *
US-AnywhereRemote in United States
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6-10+ years of experience in SRE, infrastructure, or backend systems engineering.
  • Proven ownership of reliability outcomes for complex distributed systems.
  • Strong experience with cloud infrastructure (AWS, GCP, or Azure).
  • Deep understanding of observability, incident management, and system performance.
  • Proficient in at least one programming language (e.g., Go, Python, Java).

Responsibilities

  • Lead efforts to improve system reliability, scalability, and performance across critical services.
  • Define and implement SLIs/SLOs and use them to guide engineering priorities.
  • Design and develop observability systems that produce actionable alerts.
  • Lead incident response and act as incident commander when needed.
  • Conduct postmortems to identify systemic causes and ensure corrective actions are taken.
  • Automate and eliminate operational toil through improved tooling and workflows.
  • Collaborate with product and platform teams on architecture decisions and system design.

Benefits

  • Impactful work that shapes customer experience.
  • Collaborative and inclusive culture that values big ideas and solutions.
  • Comprehensive benefits including medical, dental, vision, and wellness programs.
Full Job Description
Opportunity

We're looking for a Senior Site Reliability Engineer to help build and scale a high-impact SRE function. You'll be a technical leader on a team responsible for improving system reliability, reducing operational toil, and establishing best practices across engineering.

In this position, you'll design how reliability works in UJET, influence engineering decisions, and build the tooling and processes that make production safer and more predictable.

Responsibilities
  • Lead efforts to improve system reliability, scalability, and performance across critical services
  • Define and implement SLIs/SLOs and error budgets, and use them to guide engineering priorities
  • Design and develop observability systems (metrics, logging, tracing, alerting) that produce actionable alerts and data.
  • Lead complex incident response, acting as incident commander when needed
  • Conduct postmortems focused on systemic causes rather than individual fault, and ensure corrective actions from those reviews are completed.
  • Identify and eliminate toil through automation, tooling, and improved workflows
  • Partner with product and platform teams on architecture decisions, production readiness, and designing systems that recover from failure
  • Build reusable systems and "paved roads" that make it easier for teams to operate their services reliably
  • Mentor other engineers and raise the overall operational maturity of the organization

Requirements
  • 6-10+ years of experience in SRE, infrastructure, or backend systems engineering
  • Demonstrated experience of owning reliability outcomes for complex, distributed systems
  • Strong experience with cloud infrastructure (AWS, GCP, or Azure) and production-scale systems
  • Deep understanding of observability, incident management, and system performance
  • Proficiency in at least one programming language (e.g., Go, Python, Java) with a focus on automation and tooling
  • Able to change how other teams work without having managerial authority over them
  • Strong competency in making clear decisions during incidents by following a defined process without reacting emotionally.

Stand Out Qualifications
  • Experience building or scaling SRE practices (SLOs, incident frameworks, on-call models)
  • Kubernetes/container orchestration experience
  • Infrastructure as Code (Terraform, etc.)
  • Experience with high-growth or scaling systems
  • Background in performance engineering or capacity planning

Success Criteria
  • Critical services have clear, meaningful SLOs that drive engineering decisions
  • Alerts are actionable; irrelevant alerts are reduced; on-call workload is manageable.
  • Incidents are handled efficiently, and repeat issues decline over time
  • Engineering teams adopt reliability best practices with minimal friction
  • Toil is actively reduced through automation and better system design

Annual US Hiring Range: $140,000 - $180,000

*A candidate's actual placement within this range will depend on geographic location, work experience, education, and/or skill level.

Why UJET?
  • Impactful Work: Be at the forefront of innovation, directly shaping the future of customer experience.
  • Dynamic Culture: Join a collaborative, inclusive team that values big ideas, creative solutions, and powerful relationships.
  • Comprehensive Benefits: Medical, dental, vision, 401(k) plan, wellness benefits, and more.

Similar Jobs

More Jobs at UJET

More Information Technology Jobs

Find similar Senior Site Reliability Engineer (SRE) jobs: