Bank of Montreal

Senior Lead, Site Reliability Engineering (SRE)

Bank of Montreal • $70K — $150K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related fields.
  • 8+ years of experience in Site Reliability Engineering or related domains.
  • 3+ years of technical leadership experience for enterprise platforms.
  • Demonstrated success in driving reliability and automation initiatives across the organization.
  • Strong understanding of SRE principles, incident management, and AIOps.

Responsibilities

  • Lead the advancement of SRE practices across production platforms.
  • Establish and monitor reliability objectives such as availability and recovery metrics.
  • Develop proactive strategies to enhance reliability and customer experience.
  • Implement automation initiatives to improve service reliability and reduce operational toil.
  • Influence the adoption of engineering standards and best practices across teams.

Benefits

  • Health insurance
  • Tuition reimbursement
  • Accident and life insurance
  • Retirement savings plans
  • Opportunity to influence organizational transformation and practices.
Full Job Description

Application Deadline:

10/30/2026

Address:

4100 Gordon Baker Road

Job Family Group:

Technology

We are seeking a highly motivated and technically strong Senior Lead, Site Reliability Engineering (SRE) to provide technical leadership within the DCOE Reliability Engineering organization. This role is critical to advancing platform reliability, operational resilience, automation, and AI-driven operations across a portfolio of enterprise platforms.

The successful candidate will serve as a senior technical leader and trusted advisor, driving reliability engineering practices, modernization, automation, observability, and operational excellence across the organization. The role requires deep expertise in SRE principles, cloud and infrastructure operations, AIOps, automation engineering, service resilience, and continuous improvement.

This position is ideal for an experienced engineering leader who excels through influence, technical expertise, collaboration, and execution rather than direct people management.

Key Responsibilities

Reliability Engineering & Operational Excellence

  • Lead and advance SRE practices across critical production platforms.
  • Define, establish, and monitor platform reliability objectives, including availability, resiliency, performance, and recovery metrics.
  • Champion improvements in incident management, problem management, capacity planning, and operational readiness.
  • Develop proactive reliability strategies that reduce operational risk and improve customer experience.
  • Serve as a subject matter expert in reliability engineering, providing guidance and recommendations across technology teams.

Automation & AIOps

  • Lead automation initiatives that reduce manual effort, eliminate operational toil, and improve service reliability.
  • Design and implement automated remediation, self-healing capabilities, and operational workflows.
  • Drive adoption of AI-driven operational capabilities to enhance event correlation, anomaly detection, predictive analytics, and intelligent incident response.
  • Champion Infrastructure as Code (IaC), Configuration as Code (CaC), CI/CD integration, and automated governance controls.
  • Develop and enhance enterprise observability solutions leveraging monitoring, logging, tracing, and analytics platforms.
  • Establish standardized monitoring frameworks, operational dashboards, and reliability scorecards.
  • Drive advancements in telemetry, operational insights, and performance engineering.

Modernization & Transformation

  • Support cloud modernization, platform simplification, and technology transformation initiatives.
  • Partner with engineering, architecture, security, and operations teams to embed reliability requirements throughout the software delivery lifecycle.
  • Influence the adoption of engineering standards, best practices, and operating models across enterprise platforms.
  • Act as a key contributor to strategic technology transformation initiatives and reliability-focused roadmaps.

Technical Leadership & Engineering Enablement

  • Provide technical leadership, coaching, and mentorship to engineers and technical leads.
  • Foster a culture of accountability, innovation, continuous learning, and operational excellence.
  • Build organizational capability in automation, AI, cloud operations, observability, and reliability engineering.
  • Lead communities of practice and drive knowledge-sharing across engineering teams.
  • Influence senior stakeholders and technology partners to align on reliability, resilience, and modernization priorities.

Required Qualifications

Education

  • University degree in Computer Science, Engineering, Information Technology, or a related discipline.
  • Relevant industry certifications are considered an asset.

Experience

  • 8+ years of progressive experience in Site Reliability Engineering, Production Engineering, Infrastructure Engineering, Cloud Operations, or DevOps.
  • 3+ years of experience providing technical leadership for enterprise-scale platforms and complex engineering initiatives.
  • Proven experience leading critical production support and reliability programs through influence and collaboration.
  • Demonstrated success driving enterprise-wide reliability, automation, and modernization initiatives.

Technical Skills

  • Strong understanding of SRE principles, Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and reliability frameworks.
  • Hands-on experience with enterprise monitoring and observability platforms such as Dynatrace, Splunk, Grafana, and AppDynamics.
  • Experience implementing automation solutions using Python, PowerShell, Java, Go, or similar technologies.
  • Strong knowledge of CI/CD pipelines, GitHub, Infrastructure as Code, and cloud-native technologies.
  • Experience with AIOps tools, event management platforms, and intelligent automation solutions.
  • Knowledge of AWS, Azure, or GCP cloud platforms.
  • Experience integrating operational platforms, ITSM processes, and automated remediation capabilities.

Success Measures

The successful candidate will:

  • Improve platform reliability, resiliency, and operational stability.
  • Increase automation adoption and reduce manual operational effort.
  • Enhance incident prevention and reduction through AIOps capabilities.
  • Accelerate cloud modernization and technology transformation outcomes.
  • Strengthen engineering productivity and operational efficiency.
  • Build and elevate reliability engineering capabilities across the organization.
  • Drive measurable contributions toward BMO's Ambition 2030 objectives through innovation, accountability, automation, and engineering excellence.

Salary:

$70,000.00 - $150,000.00

Pay Type:

Salaried

The above represents BMO Financial Group’s pay range and type.

Salaries will vary based on factors such as location, skills, experience, education, and qualifications for the role, and may include a commission structure. Salaries for part-time roles will be pro-rated based on number of hours regularly worked. For commission roles, the salary listed above represents BMO Financial Group’s expected target for the first year in this position.

BMO Financial Group’s total compensation package will vary based on the pay type of the position and may include performance-based incentives, discretionary bonuses, as well as other perks and rewards. BMO also offers health insurance, tuition reimbursement, accident and life insurance, and retirement savings plans. To view more details of our benefits, please visit: 

About Bank of Montreal

The Bank of Montreal is a Canadian multinational investment bank and financial services company. It provides a wide range of personal and commercial banking, wealth management, and investment banking products and services. The bank had revenues of CAD 23.6 billion in 2020.
Learn more about Bank of Montreal
Size
45,454 employees
Market Cap
$60.9 billion
Industry
Founded
1817
5 Year Trend
+9.1%
NASDAQ

Similar Jobs

More Jobs at Bank of Montreal

More Information Technology Jobs

Find similar Senior Lead, Site Reliability Engineering (SRE) jobs: