Okta

Manager, Site Reliability Engineering (Auth0)

Okta$182K — $250K *
Consumer Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of hands-on team leadership in SRE or software engineering roles
  • 8+ years of total industry experience
  • Deep expertise in cloud platforms (AWS, Azure) and infrastructure as code (Terraform)
  • Strong programming skills in Go or Python
  • Data-driven mindset grounded in SRE principles
  • Exceptional communication skills for clarity in high-pressure situations
  • Proven ability to build and lead high-performing teams in remote-first environments

Responsibilities

  • Lead the SRE team's technical direction and driving complex, cross-functional initiatives
  • Participate in 24/7 on-call rotations to troubleshoot incidents on critical systems
  • Design and implement monitoring, alerting, and automation improvements for infrastructure resilience
  • Establish reliability best practices embedding observability into engineering efforts
  • Elevate team capabilities through mentoring, pair programming, and code reviews
  • Represent reliability in architectural reviews and strategic planning

Benefits

  • Health, dental and vision insurance
  • 401(k) and flexible spending account
  • Paid leave including PTO and parental leave
  • Immersive in-person onboarding experience
  • Equity and bonus opportunities with applicable plans
Full Job Description
The SRE Leadership Team

The SRE Leadership Team at Okta is the backbone of our platform's reliability and operational excellence. We are a forward-thinking group of engineers and leaders who believe that great infrastructure is invisible-it just works. Our team champions a culture of continuous learning, data-driven decision-making, and blameless incident response. We work at the intersection of product engineering, architecture, and operations to ensure Auth0 remains the trusted authentication platform for millions of users worldwide. As a Manager, Site Reliability Engineer, you'll lead this team with a focus on scalability, resilience, and empowering engineers to grow as technical leaders.
What You'll Be Doing
  • Lead the SRE team's technical direction, translating organizational vision into actionable roadmaps while driving complex, cross-functional initiatives across product and platform teams
  • Operate at scale through hands-on participation in 24/7 on-call rotations (follow-the-sun weekdays, shared weekends), directly troubleshooting and remediating incidents on critical systems
  • Build infrastructure resilience, designing and implementing monitoring, alerting, and automation improvements that reduce toil and elevate operational efficiency
  • Champion reliability best practices, establishing policies and cultural standards that embed observability, resilience, and software engineering rigor into all engineering efforts
  • Mentor and develop SRE talent, elevating team capabilities through pair programming, design discussions, and code reviews while fostering a culture of continuous learning
  • Represent reliability as a senior technical leader in architectural reviews and strategic planning, ensuring reliability is a core consideration in major engineering decisions
What You'll Bring to the Role
  • 3+ years of hands-on team leadership in SRE or software engineering roles within cloud-native environments, combined with 8+ years of total industry experience
  • Deep expertise in cloud platforms (AWS, Azure) and infrastructure as code (Terraform), with proven experience managing cloud-native architectures including containers, Kubernetes, microservices, and databases
  • Strong programming skills in Go or Python, with a track record of building and maintaining production-grade tools, automation, and infrastructure solutions
  • Data-driven mindset grounded in SRE principles: blameless culture, systematic problem-solving, and the ability to apply software engineering approaches to operational challenges
  • Exceptional communication skills-both verbal and written-enabling you to drive clarity during high-pressure incidents and articulate complex concepts to diverse stakeholders
  • Proven ability to build and lead high-performing teams in globally distributed, remote-first environments with strong interpersonal and collaboration skills
  • Strategic vision and technical depth, combining leadership acumen with hands-on technical excellence and a passion for mentoring senior engineers and shaping team direction
Extra Credit
  • Experience leading reliability initiatives that directly improved system uptime and reduced incident response times at scale
  • Contributions to open-source infrastructure or observability tooling
  • Experience designing and implementing comprehensive incident response programs and runbook automation

Additional requirements:
  • This position requires the ability to access federal environments and/or have access to protected federal data. As a condition of employment for this position, the successful candidate must be able to submit documentation establishing U.S. Person status (e.g. a U.S. Citizen, National, Lawful Permanent Resident, Refugee, or Asylee. 22 CFR 120.15) upon hire.

P13036

Below is the annual base salary range for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York and Washington. Your actual base salary will depend on factors such as your skills, qualifications, experience, and work location. In addition, Okta offers equity (where applicable), bonus, and benefits, including health, dental and vision insurance, 401(k), flexible spending account, and paid leave (including PTO and parental leave) in accordance with our applicable plans and policies. To learn more about our Total Rewards program please visit: https://rewards.okta.com/us.

The annual base salary range for this position for candidates located in California (excluding San Francisco Bay Area), Colorado, Illinois, New York, and Washington is between:

$182,000-$250,800 USD

The Okta Experience
  • Supporting Your Well-Being
  • Driving Social Impact
  • Developing Talent and Fostering Connection + Community

We are intentional about connection. Our global community, spanning over 20 offices worldwide, is united by a drive to innovate. Your journey begins with an immersive, in-person onboarding experience designed to accelerate your impact and connect you to our mission and team from day one.

About Okta

Okta is a leading provider of identity and access management solutions for enterprises. The company's cloud-based platform enables organizations to securely connect people and technology, providing secure access to applications and data from any device, anywhere, at any time. Okta's solutions are used by thousands of organizations worldwide, including many Fortune 500 companies. The company was founded in 2009 and is headquartered in San Francisco, California. Okta is committed to providing innovative solutions that help organizations stay secure and productive in today's digital world.
Learn more about Okta
Size
5,342 employees
Market Cap
$10.5 billion
Industry
Net Income
-$266.3 million
Founded
2009
5 Year Trend
+51.9%
Revenue
$835.4 million
NASDAQ

Similar Jobs

More Jobs at Okta

More Consumer Technology Jobs

Find similar Manager, Site Reliability Engineering (Auth0) jobs: