OC Tanner

Manager, Site Reliability Engineering

OC Tanner$120K — $150K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in Site Reliability Engineering, DevOps, Platform Engineering, or similar fields, with 2+ years in a leadership role.
  • Experience leading production operations and incident management with an emphasis on reliability engineering.
  • Proficient in designing and implementing SRE practices and operational maturity for evolving tech organizations.
  • Background in managing large-scale, customer-facing SaaS platforms emphasizing high availability and scalability.
  • In-depth knowledge of modern software engineering practices and collaboration with development teams for reliable systems.
  • Hands-on experience with observability tools like OpenTelemetry, Datadog, or Coralogix in production environments.
  • Strong AWS and Kubernetes experience, along with understanding reliability engineering principles.

Responsibilities

  • Lead and develop a Site Reliability Engineering team, promoting operational excellence and a culture of reliability.
  • Define and implement the reliability strategy to enhance availability, scalability, and performance through best practices.
  • Establish team priorities and success metrics that align with broader business goals and platform health.
  • Collaborate with Engineering, Product, and Support leaders to ensure shared ownership and optimize reliability throughout the software lifecycle.
  • Build observability capabilities and standards for metrics, logs, and Service Level Objectives (SLOs).
  • Manage incident response and post-incident reviews to improve service reliability and operational readiness.
  • Drive continuous improvement via automation and risk reduction, enhancing overall operational practices.

Benefits

  • Flexible work environment with opportunities for hybrid work.
  • Professional development and career advancement opportunities.
  • Health and wellness programs to support employee well-being and work-life balance.
  • Comprehensive healthcare benefits including medical, dental, and vision coverage.
  • Retirement savings plan with company matching contributions.
Full Job Description
Location: Salt Lake City, UT

As the Manager of Site Reliability Engineering, you will lead the strategy, execution, and evolution of reliability for our world-class employee recognition platform. You will build, mentor, and empower a team of Site Reliability Engineers while partnering closely with Engineering, Product, and Support organizations to deliver highly available, scalable, and resilient services that serve millions of users. We are seeking a leader who is passionate about operational excellence, continuous improvement, and fostering a reliability-first culture through automation, observability, and shared ownership. In this role, you will champion the development of self-healing platforms, drive incident and operational maturity, and enable engineering teams to innovate faster while delivering exceptional customer experiences.

Key Responsibilities:
  • Lead, mentor, and develop a team of Site Reliability Engineers, fostering a culture of reliability, accountability, operational excellence, and continuous improvement.
  • Define and execute the organization's reliability strategy, improving availability, scalability, performance, and resilience through automation and engineering best practices.
  • Establish team priorities, goals, and success metrics aligned with business objectives, customer needs, and platform health.
  • Partner with Engineering, Product, and Support leaders to drive shared ownership of production services and embed reliability, observability, and operational excellence throughout the software development lifecycle.
  • Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar tools, establishing enterprise standards for metrics, logs, traces, alerting, and Service Level Objectives (SLOs).
  • Oversee production triage, incident response, and escalation processes, ensuring timely service restoration, effective root cause analysis, and blameless post-incident reviews.
  • Champion a reliability-first engineering culture focused on automation, proactive risk reduction, operational readiness, shift-left quality practices, and continuous improvement.
  • Collaborate with global engineering teams in a follow-the-sun support model, ensuring seamless 24x7 coverage, effective operational handoffs, and consistent service ownership.
  • Own on-call programs, incident management practices, and operational health metrics, driving improvements in alert quality, operational efficiency, and toil reduction.
  • Manage team capacity, hiring, performance management, career development, budgeting, and workforce planning to ensure effective support of business-critical services.
  • Provide regular reporting to engineering and executive leadership on reliability trends, incidents, risks, performance metrics, and strategic initiatives.

Required Qualifications
  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related disciplines, including 2+ years in a technical leadership or people management role.
  • Proven experience leading teams responsible for production operations, reliability engineering, incident management, and operational excellence.
  • Experience designing and implementing SRE practices, reliability programs, or operational maturity initiatives within growing engineering organizations.
  • Experience operating large-scale, customer-facing SaaS platforms with high availability, performance, and scalability requirements.
  • Strong understanding of modern software engineering practices and partnering with development teams to build reliable, resilient systems.
  • Hands-on experience with observability platforms such as OpenTelemetry, Datadog, Coralogix, or similar technologies.
  • Strong knowledge of AWS and Kubernetes in production environments.
  • Deep understanding of monitoring, logging, distributed tracing, SLIs, SLOs, error budgets, and reliability engineering principles.
  • Demonstrated ability to lead cross-functional initiatives and influence stakeholders across Engineering, Product, and Support organizations.
  • Experience developing engineering roadmaps, defining team objectives, aligning reliability investments with business priorities, and driving continuous operational improvement through incident learning and post-incident reviews.


Preferred Qualifications
  • Experience leading distributed or globally dispersed engineering teams.
  • Experience with multiple cloud providers or cloud-agnostic platform architectures.
  • Familiarity with security, compliance, governance, and operational risk management frameworks.
  • Proficiency with modern Infrastructure-as-Code and technologies such as Terraform, Golang, Python, Playwright, and Performance Monitoring tools.
  • Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora.
  • Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event-driven technologies.
  • Strong understanding of cost optimization, platform sustainability, and engineering efficiency metrics.

About OC Tanner

O.C. Tanner is a global employee recognition and corporate gifting company. Founded in 1927, the company is headquartered in Salt Lake City, Utah, and has offices in Canada, the United Kingdom, and Australia. O.C. Tanner provides employee recognition programs, service awards, and incentive programs to businesses worldwide. The company has been recognized for its workplace culture and has been named one of the best places to work in the United States by the Great Place to Work Institute.
Learn more about OC Tanner
Size
2,000 employees
Industry
Founded
1927

More Jobs at OC Tanner

More Information Technology Jobs

Find similar Manager, Site Reliability Engineering jobs: