Bank of Montreal

Lead, Reliability Engineering

Bank of Montreal$121K — $211K *
Finance & Insurance
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years in tech leadership roles, preferably in financial services
  • Experience with cloud (AWS/Azure) and on-prem technologies
  • Deep expertise in observability tools, particularly Dynatrace
  • Strong incident, problem, and change management abilities
  • Proven track record in automation, with knowledge of AIOps and tools like Ansible

Responsibilities

  • Lead L2 and L3 tech support and SRE teams
  • Manage major incident processes and recovery
  • Drive service stability and resiliency across the tech stack
  • Implement AIOps and NoOps principles
  • Enable proactive operations through intelligent automation
  • Shift teams to engineering-led operations and automate workflows
  • Establish technology support strategies aligned with business priorities

Benefits

  • Health insurance
  • Tuition reimbursement
  • Life and accident insurance
  • Retirement savings plans
  • Performance-based incentives and discretionary bonuses
Full Job Description

Application Deadline:

09/03/2026

Address:

33 Dundas Street West

Job Family Group:

Technology

Transformation Mandate

This role exists to fundamentally shift Commercial Banking technology operations from reactive support to an AI-enabled, engineering-driven reliability model at scale.

The Lead will own the outcome of this transformation, challenging the status quo and driving a step-change in automation, observability, and operational excellence.

Role Overview

We are looking for a high-performing, results-driven senior technology leader to take on the role of Lead, Technology Support & Reliability Engineering. This is a critical leadership position responsible for driving the reliability, resiliency, and modernization of technology platforms supporting Commercial Banking.

The successful candidate will set the vision and lead the transformation of traditional support into a modern, engineering-led, highly automated operating model. This includes the aggressive adoption of Site Reliability Engineering (SRE), AIOps, and a forward-looking NoOps philosophy to reduce toil, increase automation, and enable self-healing, resilient systems at scale.

This role demands a proven high performer who consistently owns the outcome, questions the status quo, and delivers measurable improvements in stability, efficiency, and engineering velocity.

Key Responsibilities

Reliability & Production Support Leadership

  • Lead L2 and L3 technology support and SRE teams across Commercial Banking platforms
  • Own major incident management, escalation, and recovery processes
  • Drive service stability, availability, and resiliency outcomes through the entire tech stack, including supporting functions such as infrastructure, helpdesk, data lake, etc.

SRE, AIOps & NoOps Transformation

  • Lead the transformation toward a modern SRE-driven operating model incorporating AIOps and NoOps principles
  • Drive implementation of observability frameworks, predictive analytics, and event correlation
  • Enable proactive operations through intelligent automation and predictive capabilities
  • Reduce operational toil through automation, self-healing systems, and intelligent workflows

Engineering & Automation Enablement

  • Shift teams from reactive support to engineering-led operations
  • Build automated workflows and deployment pipelines
  • Enable end-to-end service automation and platform abstractions
  • Support cloud adoption and improve proactive operational capabilities

Strategy, Governance & Operating Model

  • Define and execute technology support strategy aligned to enterprise and Commercial Banking priorities
  • Own and evolve the L2/L3 operating model
  • Establish clear RACI and ownership models across engineering and operations
  • Drive governance, controls, and continuous improvement

Stakeholder & Delivery Leadership

  • Act as a senior technology executive and trusted advisor
  • Manage cross-functional dependencies across engineering and platform teams
  • Lead large-scale strategic initiatives and transformation programs

People Leadership & Talent Development

  • Build and lead high-performing teams with a strong culture of accountability and ownership
  • Attract, retain, and develop top engineering and SRE talent
  • Drive a performance-oriented culture aligned to enterprise objectives
  • Mentor and develop leaders capable of operating at scale

Performance Expectations & Outcomes

Reliability & Service Stability

  • Reduce Major Incident volume (P1/P2) by 30–50%
  • Improve Mean Time to Restore (MTTR) by 25–40%
  • Deliver sustained improvements in service availability and resiliency

Operational Efficiency & Automation

  • Eliminate 95% of manual operational toil
  • Increase automated incident resolution through AIOps
  • Reduce human intervention in steady-state operations

Observability, AIOps & Predictive Operations

  • Establish end-to-end observability coverage
  • Implement event correlation, predictive alerting, automated triage
  • Improve alert quality by reducing noise

Operating Model Transformation

  • Transition to SRE-led model
  • Define L2/L3 ownership
  • Shift to proactive incident prevention

Talent & Capability Development

  • Build high-performance engineering culture
  • Increase productivity and engagement
  • Develop senior SRE capability

Qualifications & Experience

Required

  • Both cloud (AWS/Azure) and on prem deployment and support
  • Deep observability experience is a must, preferably with some exposure to Dynatrace
  • Expertise in incident, problem, change management
  • Demonstrated success in automation (Ansible a plus) and AIOps

Preferred

  • Experience with Agile methodologies
  • Financial services experience
  • SRE/DevOps/platform engineering knowledge
  • ITIL Foundations or better

Leadership Profile

  • High-performance, results-oriented mindset
  • Owns the outcome and challenges status quo
  • Strategic thinker
  • Strong communicator
  • Passion for innovation and automation

Why This Role Matters

This role is central to transforming technology support into a modern, AI-enabled, highly automated reliability organization, impacting customer experience, operational risk, and delivery speed.

Salary:

$121,600.00 - $211,800.00

Pay Type:

Salaried

The above represents BMO Financial Group’s pay range and type.

Salaries will vary based on factors such as location, skills, experience, education, and qualifications for the role, and may include a commission structure. Salaries for part-time roles will be pro-rated based on number of hours regularly worked. For commission roles, the salary listed above represents BMO Financial Group’s expected target for the first year in this position.

BMO Financial Group’s total compensation package will vary based on the pay type of the position and may include performance-based incentives, discretionary bonuses, as well as other perks and rewards. BMO also offers health insurance, tuition reimbursement, accident and life insurance, and retirement savings plans. To view more details of our benefits, please visit: 

About Bank of Montreal

The Bank of Montreal is a Canadian multinational investment bank and financial services company. It provides a wide range of personal and commercial banking, wealth management, and investment banking products and services. The bank had revenues of CAD 23.6 billion in 2020.
Learn more about Bank of Montreal
Size
45,454 employees
Market Cap
$60.9 billion
Industry
Founded
1817
5 Year Trend
+9.1%
NASDAQ

Similar Jobs

More Jobs at Bank of Montreal

More Finance & Insurance Jobs

Find similar Lead, Reliability Engineering jobs: