Sr. Manager, Site Reliability Engineer

iCIMS Talent Acquisition

$150K — $170K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years in Site Reliability Engineering, Cloud Engineering, or related fields.
  • 5+ years in a technical leadership role with strategic and organizational responsibilities.
  • Experience managing distributed teams across various time zones.
  • Proficient in AWS cloud architecture and complex enterprise environments.
  • Broad knowledge of cloud computing technologies including containers and networking.
  • Strong grasp of scalable and resilient production systems.
  • Familiarity with Infrastructure as Code, ideally using Terraform.

Responsibilities

  • Lead and grow a global Site Reliability Engineering team across multiple regions.
  • Define the SRE strategy and set technical directions.
  • Establish a hybrid SRE model for centralized support and product alignment.
  • Clarify roles and authority within SRE and engineering teams.
  • Develop talent through mentoring and clear expectations.
  • Cultivate a proactive reliability engineering culture.
  • Collaborate with product teams to identify reliability risks and operational needs.

Benefits

  • Comprehensive health and wellness plans including medical, dental, and vision.
  • 401(k) with company match and dependent care assistance.
  • Short and long-term disability insurance, life and AD&D coverage.
  • Parental leave policies and mindfulness resources.
  • Open vacation policy, sick days, and paid holidays.
  • Quiet hours to promote well-being during workdays.
  • Tuition reimbursement for continued education opportunities.
Full Job Description
Job Overview

We are seeking a Sr Manager of Site Reliability Engineering (SRE) to lead and continue developing our global SRE organization across the US, Ireland, India, and strategic engineering partners.

This role is responsible for advancing a modern SRE operating model that combines centralized reliability capabilities with product-aligned SRE support. The Sr Manager will partner closely with Product Engineering, Cloud Engineering, Operations, Security, DBA, and other technical teams to improve reliability, strengthen operational practices, and help development teams implement consistent engineering standards.

This is a highly technical leadership role requiring strong experience across Site Reliability Engineering, AWS cloud architecture, observability, cloud engineering, automation, incident and problem management, and FinOps. The successful candidate will be able to evaluate technical decisions across reliability, scalability, security, operational complexity, and cloud financial impact.

Responsibilities

Leadership & Strategy
  • Lead and develop a globally distributed SRE organization across multiple regions and time zones.
  • Define and execute the SRE strategy, operating model, priorities, and technical direction.
  • Establish a hybrid SRE model combining centralized Reliability Enablement capabilities with product-aligned SRE support.
  • Define clear responsibilities and decision rights across SRE, Product Engineering, Operations, Cloud Engineering, DBA, and other technical functions.
  • Build strong technical leadership, product ownership, regional handoffs, and knowledge-sharing practices across the global organization.
  • Develop engineers and technical leaders through coaching, mentorship, and clear technical and career expectations.
  • Drive a culture focused on proactive reliability engineering rather than reactive operational support.


Reliability Engineering & Product Alignment
  • Partner with Product Engineering teams to understand product architecture, service dependencies, reliability risks, and operational requirements.
  • Establish and mature SLIs, SLOs, error budgets, service-health measures, and operational-readiness standards where appropriate.
  • Identify systemic reliability issues before they become customer-impacting incidents.
  • Translate production learnings, recurring failures, and RCAs into prioritized engineering improvements.
  • Establish reusable reliability patterns, tooling, automation, runbooks, and engineering practices that can be adopted across product teams.
  • Ensure SRE supports product teams without replacing Product Engineering ownership of application functionality and defects.


Problem Management & Operational Excellence
  • Provide technical leadership during significant and complex production incidents.
  • Partner with Operations and engineering teams to strengthen incident response, escalation, restoration, and recovery practices.
  • Lead the continued development of structured problem management and root-cause analysis practices.
  • Connect recurring incidents and operational risks to visible, prioritized corrective actions.
  • Improve operational readiness, capacity planning, performance management, and service resiliency.
  • Reduce repetitive operational work and manual intervention through automation and engineering.


Observability & Reliability Enablement
  • Lead the development and adoption of common observability standards across logging, metrics, tracing, dashboards, monitoring, and alerting.
  • Drive enterprise observability strategy and governance across platforms including Grafana, OpenTelemetry, Sumo Logic, New Relic, CloudWatch, and related technologies.
  • Establish practical service blueprints and reusable observability patterns for development teams.
  • Improve alert quality, service visibility, dependency awareness, and actionable monitoring.
  • Ensure observability capabilities support both real-time incident response and longer-term reliability improvement.


Operational FinOps & Cloud Financial Management
  • Incorporate cloud financial awareness into architecture, reliability, and engineering decisions.
  • Partner with Cloud Engineering, Finance, Product, and engineering leadership to improve cloud cost ownership and accountability.
  • Identify opportunities for rightsizing, workload optimization, storage efficiency, and removal of unnecessary cloud consumption.
  • Understand AWS pricing and commitment constructs including Savings Plans, Reserved Instances, licensing considerations, and consumption-based services.
  • Support effective cost allocation, tagging, forecasting, reporting, and cloud financial governance.
  • Evaluate technical decisions using both engineering and financial considerations while ensuring cost optimization does not introduce unacceptable reliability or performance risk.


AWS & Cloud Engineering
  • Provide senior technical leadership for complex cloud environments, with AWS as the primary platform.
  • Partner with Cloud Engineering on architecture, resiliency, automation, networking, security, governance, and platform standards.
  • Review and guide architectures involving AWS technologies such as ECS, ECR, EC2, RDS, Aurora, S3, DynamoDB, OpenSearch, SQS, SNS, Kinesis, IAM, AWS Organizations, and cloud networking.
  • Evaluate architecture across availability, scalability, performance, security, recoverability, operational complexity, and cost.
  • Drive Infrastructure as Code, automated provisioning, standardized cloud patterns, and policy-based governance

Qualifications

  • 10+ years of experience across Site Reliability Engineering, Cloud Engineering, Platform Engineering, DevOps, Infrastructure Engineering, or related technical disciplines.
  • 5+ years of technical or engineering leadership experience, including responsibility for technical strategy, team leadership, and organizational outcomes.
  • Experience leading and collaborating with distributed technical teams across multiple regions and time zones.
  • Strong technical knowledge of AWS and experience supporting or designing complex enterprise cloud environments.
  • Broad understanding of cloud computing, containers, networking, storage, databases, security, identity, monitoring, and governance.
  • Strong understanding of highly available, scalable, resilient, and distributed production systems.
  • Experience with Infrastructure as Code and automated infrastructure delivery, preferably Terraform.
  • Strong understanding of modern observability practices including logs, metrics, traces, monitoring, dashboards, and alerting.
  • Demonstrated experience with incident response, technical escalation, root-cause analysis, and problem management.
  • Strong understanding of FinOps and cloud financial management principles, including optimization, allocation, forecasting, and cost accountability.
  • Ability to evaluate architecture from both technical and financial perspectives.
  • Strong communication and collaboration skills with the ability to influence engineers, architects, Product leaders, Finance, Security, and executive stakeholders.

Compensation and Benefits

We accept applications for this position on an ongoing basis until the position is filled. Applications will be reviewed as they are received, and qualified candidates may be contacted throughout the posting period.

The anticipated base salary range for this position is $150,000 - $170,000. In addition, the estimated on-target earnings ("OTE"), which includes base salary and commissions, is $170,000-$200,000.

Actual compensation will depend on various job-related factors, including but not limited to, location, experience, and job qualifications. This range aligns with our commitment to equitable and transparent compensation practices, as required by applicable law.

Competitive health and wellness benefits include medical, dental, vision, 401(k), dependent care, short term and long-term disability, life and AD&D insurance, bonding and parental leave, mindfulness resources, an open vacation policy, sick days, paid holidays, quiet hours each workday, and tuition reimbursement. Benefits and eligibility may vary by location, role, and tenure. Learn more here: https://careers.icims.com/benefits

Similar Jobs

More Jobs at iCIMS Talent Acquisition

More Information Technology Jobs

Find similar Sr. Manager, Site Reliability Engineer jobs: