Manager, Site Reliability Engineering62 M Overview Responsible for leading the Site Reliability Engineering Center of Excellence and the Forward Deployed SRE program supporting critical banking platforms, applications, and technology services. Manages an organization of employees and contingent resources through direct reports, program managers, SRE leaders, technical leads, and matrixed delivery relationships.
Establishes the enterprise SRE strategy, operating model, engineering standards, governance, talent model, and adoption roadmap. Accountable for improving service reliability, availability, scalability, performance, resiliency, deployment safety, and operational maturity across supported technology domains.
Leads the deployment of SRE capabilities into application and platform teams through a Forward Deployed SRE model. Partners with senior leaders across application development, infrastructure, cloud engineering, architecture, cybersecurity, technology operations, risk, and business-aligned technology organizations to prioritize engagements and deliver measurable reliability improvements.
Balances strategic leadership, people management, program execution, technical governance, and operational accountability. Ensures SRE practices are implemented consistently and that reliability investments produce measurable improvements in customer experience, operational risk, engineering productivity, release quality, and service performance.
Primary Responsibilities SRE Strategy and Center of Excellence Leadership - Establish and execute the vision, strategy, operating model, service offerings, and multiyear roadmap for the Site Reliability Engineering Center of Excellence.
- Define enterprise SRE standards, engineering practices, governance processes, engagement models, and maturity expectations.
- Lead the adoption of reliability engineering practices across application development, infrastructure, platform engineering, cloud engineering, and technology operations.
- Translate enterprise technology, business, resiliency, and risk priorities into an actionable SRE portfolio and delivery roadmap.
- Establish a scalable SRE service model that includes consulting, enablement, embedded engineering, Forward Deployed SRE engagements, reusable capabilities, and sustained ownership by application and platform teams.
- Define intake, prioritization, engagement, transition, and exit criteria for SRE services.
- Develop and maintain an SRE maturity model used to assess service teams, identify reliability gaps, and guide improvement plans.
- Establish communities of practice, technical forums, training programs, playbooks, reference architectures, and reusable engineering patterns that expand SRE capabilities across the organization.
- Ensure the SRE Center of Excellence remains aligned with enterprise engineering standards, cloud strategy, operational risk requirements, and evolving industry practices.
- Represent the SRE organization in senior leadership forums, architecture reviews, operational governance meetings, and enterprise transformation initiatives.
Forward Deployed SRE Program Leadership - Lead the Forward Deployed SRE program, placing SRE professionals into high-priority application and platform teams to address complex reliability challenges and improve operational maturity.
- Manage program managers, SRE leaders, and technical leads responsible for coordinating engagements across multiple technology domains.
- Establish a transparent intake and prioritization process based on customer impact, service criticality, operational risk, incident history, reliability maturity, strategic importance, and anticipated business value.
- Partner with application and platform leaders to define engagement objectives, scope, deliverables, staffing, success measures, dependencies, and duration.
- Ensure Forward Deployed SRE teams deliver sustainable engineering improvements rather than becoming long-term substitutes for application support or production operations.
- Establish shared accountability for participation, knowledge transfer, remediation activities, and long-term ownership of implemented reliability capabilities.
- Develop transition and exit plans that enable application and platform teams to sustain SRE practices after engagements conclude.
- Evaluate engagement effectiveness using measurable outcomes such as availability, SLO attainment, incident frequency, restoration time, alert quality, automation adoption, toil reduction, change-failure rate, deployment reliability, and engineering maturity.
- Convert common findings and lessons learned into reusable standards, automation, tools, training, and engineering patterns.
- Continuously optimize the Forward Deployed SRE operating model based on demand, capacity, outcomes, stakeholder feedback, and changes in technology strategy.
Organizational and People Leadership - Lead an organization of employees and contingent resources through direct and indirect management relationships.
- Manage and develop program managers, SRE managers, technical leaders, and senior engineering professionals.
- Establish clear roles, responsibilities, decision rights, performance expectations, and accountability across the SRE organization.
- Build a high-performing organization with expertise in reliability engineering, observability, cloud platforms, release engineering, deployment orchestration, automation, Infrastructure as Code, incident management, resiliency engineering, testing, and program delivery.
- Recruit, retain, coach, and develop diverse engineering and program management talent.
- Conduct workforce, capacity, and succession planning to ensure the organization has the leadership and technical capabilities required to meet current and future demand.
- Define career paths and skill-development plans for SRE professionals in partnership with engineering and human resources leaders.
- Provide regular feedback, performance management, recognition, coaching, and development opportunities.
- Create an environment that encourages engineering discipline, continuous learning, constructive challenge, accountability, collaboration, and knowledge sharing.
- Promote an environment that supports a culture of belonging and reflects the M&T Bank brand.
- Manage staffing and sourcing strategies, including the appropriate use of employees, contingent labor, managed services, and specialized partners.
- Ensure third-party resources and service providers meet applicable engineering, security, risk, performance, financial, and contractual expectations.
Reliability Engineering Governance - Establish standards and governance for Service Level Indicators, Service Level Objectives, error budgets, availability targets, and reliability reporting.
- Partner with service owners and business stakeholders to align reliability objectives with customer expectations, service criticality, business impact, risk tolerance, and cost.
- Define common reliability metrics and executive-level reporting for service health, operational performance, deployment health, risk, and improvement outcomes.
- Ensure reliability metrics are accurate, actionable, consistently defined, and tied to accountable service owners.
- Establish governance for error-budget decisions, including remediation, release-risk evaluation, delivery tradeoffs, escalation, and investment prioritization.
- Drive adoption of reliability-by-design practices throughout the Software Development Lifecycle.
- Establish operational readiness requirements for applications and platforms entering production or undergoing material change.
- Ensure readiness assessments address availability, scalability, observability, performance, supportability, recoverability, security, testing, capacity, release and rollback readiness, documentation, and operational ownership.
- Partner with architecture and engineering leaders to incorporate reliability requirements into solution designs, architecture reviews, and engineering standards.
- Identify systemic reliability risks and sponsor cross-organizational remediation programs.
Release Engineering and Deployment Orchestration - Establish enterprise SRE standards for safe, repeatable, observable, and recoverable application and infrastructure releases.
- Partner with application development, platform engineering, cloud engineering, quality engineering, change management, and technology operations teams to improve release and deployment practices.
- Define approved deployment patterns based on service criticality, architecture, customer impact, technical capability, and risk.
- Promote progressive delivery practices, including blue-green deployments, canary releases, rolling deployments, ring-based deployments, feature flags, traffic splitting, and controlled production experimentation where appropriate.
- Establish requirements for automated pre-deployment, in-deployment, and post-deployment validation using technical health signals, service-level indicators, business metrics, and customer-experience measures.
- Promote deployment orchestration that integrates CI/CD pipelines, Infrastructure as Code, automated testing, observability, approval controls, and policy enforcement.
- Establish standards for automated rollback, roll-forward, traffic evacuation, deployment pausing, and recovery when release health thresholds are breached.
- Define release health criteria and quality gates based on error rates, latency, saturation, availability, dependency health, business transactions, and customer-impact indicators.
- Promote small, incremental, independently deployable changes that reduce release complexity and limit failure impact.
- Partner with platform teams to provide reusable deployment templates, pipeline capabilities, policy-as-code controls, and self-service release patterns.
- Establish traceability between changes, deployments, configuration updates, incidents, service telemetry, and business outcomes.
- Improve release observability through deployment markers, version-aware dashboards, automated change correlation, and real-time health analysis.
- Use error budgets, service criticality, testing evidence, deployment history, and current service health to inform release-risk decisions.
- Measure and improve deployment frequency, lead time for changes, change-failure rate, rollback effectiveness, failed deployment recovery time, and release-related customer impact.
- Ensure deployment practices comply with applicable technology risk, cybersecurity, change-management, segregation-of-duties, and regulatory requirements.
Observability, Automation, and Cloud Engineering - Define the enterprise observability strategy for supported services, including standards for telemetry, distributed tracing, metrics, logging, dashboards, alerting, dependency mapping, and customer-experience monitoring.
- Govern the effective implementation and use of Dynatrace, OpenTelemetry, Azure Monitor, Application Insights, Log Analytics, and other approved enterprise tools.
- Drive standardization of telemetry and observability patterns across cloud-native, hybrid, distributed, and legacy environments.
- Establish expectations for actionable alerts, signal quality, service-health visibility, release observability, and real-time operational insight.
- Reduce alert fatigue and operational noise by improving monitoring coverage, alert thresholds, routing, correlation, suppression, ownership, and automation.
- Sponsor reusable dashboards, instrumentation patterns, automation libraries, reference architectures, and reliability controls.
- Develop and execute a strategy to reduce operational toil through automation, self-service capabilities, automated recovery, and self-healing solutions.
- Establish measurable toil-reduction goals and ensure reclaimed engineering capacity is redirected toward reliability, automation, and product improvements.
- Promote Infrastructure as Code using Terraform and other approved technologies to improve repeatability, control, recoverability, and environment consistency.
- Partner with Azure and enterprise platform teams to improve scalability, resiliency, deployment safety, lifecycle management, and operational controls.
Incident, Problem, and Resiliency Management - Provide leadership and executive coordination during significant technology incidents affecting customers, critical business services, or enterprise operations.
- Ensure appropriate technical leadership, stakeholder communication, escalation, decision-making, and recovery focus during high-severity events.
- Partner with incident management, technology operations, application teams, infrastructure teams, and business leaders to minimize customer impact and restore services safely.
- Establish expectations for timely, objective, and technically rigorous Root Cause Analysis.
- Ensure corrective and preventive actions address underlying technical, process, monitoring, testing, release, deployment, architecture, and organizational causes.
- Track material reliability actions to completion and escalate overdue or inadequately addressed risks.