Bank of Montreal

Sr. SRE Engineer (CCaaS)

Bank of Montreal$70K — $150K *
US-AnywhereRemote in Ontario, CA
Telecommunications & Hardware
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5-6 years of relevant experience in Site Reliability Engineering or a related field.
  • Post-secondary degree in a related field or equivalent experience.
  • Advanced proficiency in configuring observability tools like Dynatrace or CloudWatch.
  • Deep knowledge of automation practices and pipelines.
  • Intermediate experience with cloud computing and container orchestration.

Responsibilities

  • Ensure 24x7 availability of the CCaaS platform and services.
  • Build and maintain monitoring dashboards, alerts, and health checks.
  • Drive incident response efforts and lead root cause analysis.
  • Automate operational processes to increase efficiency.
  • Support platform configurations, disaster recovery, and backup testing.

Benefits

  • Health insurance coverage.
  • Tuition reimbursement for continued learning.
  • Accident and life insurance options.
  • Retirement savings plans for future security.
  • Performance-based incentives and discretionary bonuses.
Full Job Description
Application Deadline:

09/16/2026

Address:
VIRTUAL59 - REMOTE/TELETRAVAIL - ON - BMO

Job Family Group:

Technology

Key Responsibilities
Platform Reliability & Availability
  • Ensure 24x7 availability of the CCaaS platform and supporting services.
  • Define and manage Service Level Indicators (SLI), Service Level Objectives (SLO), and Error Budgets.
  • Monitor platform health, performance, capacity, and stability.
  • Drive reliability improvements and eliminate recurring operational issues.
  • Support production readiness reviews for new releases and platform enhancements.
Monitoring & Observability
  • Build and maintain monitoring dashboards, alerts, and health checks.
  • Configure and optimize CloudWatch, Splunk, Dynatrace, Grafana, and other observability tools.
  • Establish proactive alerting and predictive monitoring capabilities.
  • Develop operational runbooks and troubleshooting guides.
Incident & Problem Management
  • Act as an escalation point for critical incidents and service disruptions.
  • Lead incident triage, impact assessment, communication, and recovery activities.
  • Perform Root Cause Analysis (RCA) and identify preventive actions.
  • Drive reduction of MTTR, MTTD, and recurring incidents.
  • Participate in command center and major incident bridge calls during production events.
Automation & Engineering Excellence
  • Eliminate operational toil through automation.
  • Develop self-healing capabilities and automated remediation solutions.
  • Build scripts and tools using Python, PowerShell, AWS Lambda, or other automation frameworks.
  • Automate operational checks, deployment validation, monitoring, and reporting.
Cloud & Platform Operations
  • Support Amazon Connect, AWS services, APIs, middleware, routing, IVR, and integration components.
  • Manage platform configurations, certificates, service accounts, and connectivity requirements.
  • Support disaster recovery, backup, restoration, and failover testing activities.
  • Conduct capacity planning and performance optimization.
Release & Change Support
  • Participate in release planning, deployment validation, and production implementation activities.
  • Review changes for operational risk and reliability impact.
  • Support production deployments and rollback planning.
  • Ensure operational readiness requirements are met before launch.
Security & Compliance
  • Ensure adherence to enterprise security standards and regulatory controls.
  • Support vulnerability remediation and audit activities.
  • Monitor platform security events and compliance requirements.
  • Maintain operational controls and evidence for risk reviews.
Continuous Improvement
  • Analyze operational metrics and identify improvement opportunities.
  • Drive reliability engineering best practices across CCaaS teams.
  • Participate in architecture reviews with a focus on scalability and resiliency.
  • Mentor developers and operations teams on SRE practices.


The CCaaS Site Reliability Engineer (SRE) is responsible for ensuring the reliability, availability, performance, security, and operational excellence of the Contact Centre as a Service (CCaaS) platform. This role supports Amazon Connect, IVR, routing, APIs, Lambda functions, reporting pipelines, and integrated third-party platforms such as Verint, Nuance, Pindrop, NICE, TPSS, and CRM systems.

The SRE will drive automation, observability, incident response, resiliency engineering, and continuous service improvement while partnering with Development, DevOps, Infrastructure, Security, Business, and Vendor teams.

  • Deploys, configures, and monitors code as well as the availability, latency, change management, emergency response, and management capacity of services in production.
  • Helps the development and operations teams establish Service level indicators (SLIs), Service level objectives (SLOs) and Error budgets.
  • Performs automation to increase efficiency and decrease risk like log analysis, performance tuning, patch application, testing of production settings, incident response, and post-mortem analysis.
  • Supports in system design consulting, platform management, and capacity planning.
  • Debugs production issues across services and levels of the technology stack.
  • Improves service health visibility by recording metrics, logs, and traces across all services in order to pinpoint the reasons of an incident.
  • Computes the cost of SLA breaches and assists management in calculating the impact of system reliability. Helps development and operations teams understand the cost of downtime.
  • Focus is primarily on business/group within BMO; may have broader, enterprise-wide focus.
  • Provides specialized consulting, analytical and technical support.
  • Exercises judgment to identify, diagnose, and solve problems within given rules.
  • Works independently and regularly handles non-routine situations.
  • Broader work or accountabilities may be assigned as needed.
  • Take measured risks while protecting the bank by applying our Risk Management Framework in the execution of your role, in line with our Risk Culture and within our approved Risk Appetite, making sound and risk informed decisions that align to business strategy, protect assets, and adhere to applicable policy documents (Frameworks, Policies, Standards, Procedures and Supporting documents), laws and regulations.


Qualifications:

Foundational level of proficiency:
  • DevOps.
  • Cybersecurity and privacy concepts, principles and solutions.
  • Emotional agility

Intermediate level of proficiency:
  • IT infrastructure library.
  • Robot Process Automation.
  • Cloud Computing.
  • Configuration Management.
  • Container Orchestration.
  • System Design and Implementation.
  • Incident management.
  • Learning Agility.
  • Building and managing relationships.
  • Verbal & written communication skills.
  • Collaboration & team skills.
  • Analytical and problem solving skills.
  • Data driven decision making.

Advanced level of proficiency:
  • alerting & log configuration with opensearch, cloud watch
  • Dynatrace or any other APM tool configuration & dashboarding
  • Automation and Automation Pipelines.
  • Automated Testing.
  • Typically between 5 - 6 years of relevant experience and post-secondary degree in related field of study or an equivalent combination of education and experience.
  • Deep knowledge and technical proficiency gained through extensive education and business experience.


Salary:

$70,000.00 - $150,000.00

Pay Type:

Salaried

The above represents BMO Financial Group's pay range and type.

Salaries will vary based on factors such as location, skills, experience, education, and qualifications for the role, and may include a commission structure. Salaries for part-time roles will be pro-rated based on number of hours regularly worked. For commission roles, the salary listed above represents BMO Financial Group's expected target for the first year in this position.

BMO Financial Group's total compensation package will vary based on the pay type of the position and may include performance-based incentives, discretionary bonuses, as well as other perks and rewards. BMO also offers health insurance, tuition reimbursement, accident and life insurance, and retirement savings plans. To view more details of our benefits, please visit: https://jobs.bmo.com/global/en/Total-Rewards

About Bank of Montreal

The Bank of Montreal is a Canadian multinational investment bank and financial services company. It provides a wide range of personal and commercial banking, wealth management, and investment banking products and services. The bank had revenues of CAD 23.6 billion in 2020.
Learn more about Bank of Montreal
Size
45,454 employees
Market Cap
$60.9 billion
Industry
Founded
1817
5 Year Trend
+9.1%
NASDAQ

Similar Jobs

More Jobs at Bank of Montreal

More Telecommunications & Hardware Jobs

Find similar Sr. SRE Engineer (CCaaS) jobs: