Manager, Site Reliability Engineering

LayerZero Labs

$125K — $150K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science or equivalent experience
  • 6+ years in SRE, DevOps, or infrastructure engineering
  • 2+ years in a management or leadership role
  • Deep understanding of blockchain node infrastructure
  • Strong skills in TypeScript or Golang
  • Experience running Kubernetes in production for 3+ years
  • Proven track record in enhancing on-call and incident response processes

Responsibilities

  • Lead and develop a team of SREs
  • Own the reliability strategy for blockchain node infrastructure
  • Partner with Engineering and Product teams to align reliability with business goals
  • Drive infrastructure-as-code practices using Kubernetes and Helm
  • Establish an improved on-call structure and incident management processes
  • Engage directly in reviewing designs and complex incident resolutions

Benefits

  • Comprehensive healthcare coverage
  • Flexible work arrangements
  • Professional development opportunities
  • Innovative work culture focused on technology
  • Exposure to cutting-edge blockchain technologies
Full Job Description
ABOUT THE ROLE

At LayerZero, our Site Reliability Engineering (SRE) team is at the intersection of software and systems engineering, dedicated to crafting and maintaining large-scale, resilient systems. Our goal is to ensure that all LayerZero services - ranging from critical internal systems to those external users interact with - are reliable, meet the uptime expectations of our users, and continuously evolve at a swift pace. Our SRE professionals will monitor our system's capacity and performance to uphold these standards.

As Manager of SRE, you'll lead a team of engineers responsible for the reliability, performance, and scalability of our blockchain node infrastructure and platform services - while staying technically sharp enough to guide architecture decisions and jump into critical incidents. You'll balance people leadership with hands-on technical judgment, shaping how the team works, grows, and scales alongside LayerZero.

WHAT YOU'LL DO
  • Lead and develop a team of SREs - setting technical direction, growth plans, and performance expectations.
  • Own the reliability strategy for blockchain node infrastructure across a variety of DLTs, including SLOs, capacity planning, and incident response.
  • Partner with Engineering leadership and Product/Platform teams to align reliability investments with business priorities.
  • Drive infrastructure-as-code practices, with a focus on Kubernetes and Helm at scale.
  • Establish and continuously improve on-call structure, incident detection/triage automation, and postmortem culture.
  • Stay hands-on: review designs, dig into complex incidents, and set the technical bar for the team.


ABOUT YOU
  • Bachelor's degree in Computer Science, similar technical field of study, or equivalent practical experience.
  • 6+ years in SRE, DevOps, or infrastructure engineering, including 2+ years directly managing or leading a technical team.
  • Deep familiarity with blockchain node infrastructure (validator/full/archive nodes, RPC optimization, etc.).
  • Strong proficiency in TypeScript or Golang, with the judgment to know when to write code vs. delegate.
  • Advanced knowledge of Unix/Linux internals and distributed systems / high-availability design.
  • 3+ years running Kubernetes in production, including Helm chart authoring at scale.
  • Track record building or scaling an on-call/incident response process.
  • Excellent communication skills - able to represent the team to leadership and hire/retain strong engineers.


Similar Jobs

More Jobs at LayerZero Labs

More Enterprise Technology Jobs

Find similar Manager, Site Reliability Engineering jobs: