Location: Salt Lake City, UT
As the Manager of Site Reliability Engineering, you will lead the strategy, execution, and evolution of reliability for our world-class employee recognition platform. You will build, mentor, and empower a team of Site Reliability Engineers while partnering closely with Engineering, Product, and Support organizations to deliver highly available, scalable, and resilient services that serve millions of users. We are seeking a leader who is passionate about operational excellence, continuous improvement, and fostering a reliability-first culture through automation, observability, and shared ownership. In this role, you will champion the development of self-healing platforms, drive incident and operational maturity, and enable engineering teams to innovate faster while delivering exceptional customer experiences.
Key Responsibilities:
- Lead, mentor, and develop a team of Site Reliability Engineers, fostering a culture of reliability, accountability, operational excellence, and continuous improvement.
- Define and execute the organization's reliability strategy, improving availability, scalability, performance, and resilience through automation and engineering best practices.
- Establish team priorities, goals, and success metrics aligned with business objectives, customer needs, and platform health.
- Partner with Engineering, Product, and Support leaders to drive shared ownership of production services and embed reliability, observability, and operational excellence throughout the software development lifecycle.
- Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar tools, establishing enterprise standards for metrics, logs, traces, alerting, and Service Level Objectives (SLOs).
- Oversee production triage, incident response, and escalation processes, ensuring timely service restoration, effective root cause analysis, and blameless post-incident reviews.
- Champion a reliability-first engineering culture focused on automation, proactive risk reduction, operational readiness, shift-left quality practices, and continuous improvement.
- Collaborate with global engineering teams in a follow-the-sun support model, ensuring seamless 24x7 coverage, effective operational handoffs, and consistent service ownership.
- Own on-call programs, incident management practices, and operational health metrics, driving improvements in alert quality, operational efficiency, and toil reduction.
- Manage team capacity, hiring, performance management, career development, budgeting, and workforce planning to ensure effective support of business-critical services.
- Provide regular reporting to engineering and executive leadership on reliability trends, incidents, risks, performance metrics, and strategic initiatives.
Required Qualifications- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related disciplines, including 2+ years in a technical leadership or people management role.
- Proven experience leading teams responsible for production operations, reliability engineering, incident management, and operational excellence.
- Experience designing and implementing SRE practices, reliability programs, or operational maturity initiatives within growing engineering organizations.
- Experience operating large-scale, customer-facing SaaS platforms with high availability, performance, and scalability requirements.
- Strong understanding of modern software engineering practices and partnering with development teams to build reliable, resilient systems.
- Hands-on experience with observability platforms such as OpenTelemetry, Datadog, Coralogix, or similar technologies.
- Strong knowledge of AWS and Kubernetes in production environments.
- Deep understanding of monitoring, logging, distributed tracing, SLIs, SLOs, error budgets, and reliability engineering principles.
- Demonstrated ability to lead cross-functional initiatives and influence stakeholders across Engineering, Product, and Support organizations.
- Experience developing engineering roadmaps, defining team objectives, aligning reliability investments with business priorities, and driving continuous operational improvement through incident learning and post-incident reviews.
Preferred Qualifications- Experience leading distributed or globally dispersed engineering teams.
- Experience with multiple cloud providers or cloud-agnostic platform architectures.
- Familiarity with security, compliance, governance, and operational risk management frameworks.
- Proficiency with modern Infrastructure-as-Code and technologies such as Terraform, Golang, Python, Playwright, and Performance Monitoring tools.
- Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora.
- Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event-driven technologies.
- Strong understanding of cost optimization, platform sustainability, and engineering efficiency metrics.