Royal Bank of Canada

Site Reliability Engineer (SRE), Cloud Operations

Royal Bank of Canada$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in Site Reliability Engineering, DevOps, or infrastructure operations.
  • Strong knowledge of Kubernetes/OpenShift administration and troubleshooting.
  • Hands-on experience with Ansible and Terraform for Infrastructure as Code.
  • Proficiency in Python scripting for automation and platform development.
  • Experience with monitoring and observability tools like Prometheus and Grafana.
  • Familiarity with incident management processes and on-call rotations.
  • Understanding of security and compliance fundamentals.

Responsibilities

  • Support scalable and secure architectures across cloud platforms.
  • Automate infrastructure workflows to minimize manual work.
  • Extend self-healing automation capabilities using Ansible.
  • Lead design reviews for new platform features and operational integration.
  • Collaborate with platform teams for technical feedback and code contributions.
  • Drive CI/CD and Infrastructure as Code practices using Ansible and Terraform.
  • Participate in on-call rotation for incident management and troubleshooting.

Benefits

  • Comprehensive Total Rewards Program with bonuses and flexible benefits.
  • Leaders committed to supporting your development through coaching.
  • Opportunity to make a meaningful impact within the organization.
  • Dynamic and high-performing team environment.
  • Flexible work/life balance options.
Full Job Description
Job Description

What is the opportunity?

Join the Platform Engineering & AI Operations team within OTK0, where you'll sit at the intersection of Site Reliability Engineering and intelligent infrastructure operations. This role offers the chance to shape how the bank operates, monitors, and self-heals its private and public cloud platforms - from OpenShift clusters and Kafka environments to self-healing automation systems. You'll work on real problems at enterprise scale: reducing toil for NOC, Data Center, and Branch teams, building automation that eliminates manual work, and establishing reliable operational practices. If you want to move beyond traditional ops into the future of intelligent, autonomous infrastructure operations - this is the role.

What will you do?
  • Support highly scalable, secure, and highly available architectures across private and public cloud platforms (Kubernetes/OpenShift, ECE, Confluent Kafka).
  • Write code and scripts to automate infrastructure workflows and eliminate toil, including automation pipelines that reduce manual intervention across Data Center, Branch, and NOC operations.
  • Extend self-healing automation capabilities built on Ansible, automating routine operational tasks (e.g., CPU remediation) to reduce manual intervention.
  • Participate in and lead design reviews for new platform features, infrastructure changes, and operational integration points, ensuring alignment with security, reliability, and regulatory requirements.
  • Collaborate with platform teams to provide technical feedback, contribute code changes to shared repositories, and establish data standards and pipelines (e.g., ServiceNow, Prometheus) that support operational excellence.
  • Drive automation, CI/CD, and Infrastructure as Code practices across the team, leveraging Ansible and Terraform for deployment validation and self-healing remediation workflows.
  • Minimize risk of reliability failures related to durability, availability, performance, and correctness, leveraging proactive alerting and anomaly detection.
  • Participate in on-call rotation for platform support, incident management, and troubleshooting, triaging incidents via Grafana, Prometheus, Dynatrace, and PagerDuty.


What do you need to succeed?

Must-have
  • 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or infrastructure operations.
  • Strong working knowledge of Kubernetes/OpenShift administration and troubleshooting in enterprise environments.
  • Hands-on experience with Ansible and Terraform for Infrastructure as Code and automation.
  • Proficiency in Python scripting (core to infrastructure automation and platform development).
  • Hands-on experience with monitoring and observability stacks (Prometheus, Grafana, ELK, or equivalent).
  • Experience with incident management processes, on-call rotations, and post-incident review practices.
  • Familiarity with capacity planning, threshold-based alerting, and performance trend analysis.
  • Understanding of security and compliance fundamentals, including vulnerability assessment and remediation tracking.
  • Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).


Nice-to-have
  • Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).
  • Hands-on experience with public cloud platforms (AWS, Azure, GCP) in hybrid or multi-cloud environments.
  • Experience with GPU/compute infrastructure for ML inference workloads


What's in it for you?

We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual.
  • A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable
  • Leaders who support your development through coaching and managing opportunities
  • Ability to make a difference and lasting impact
  • Work in a dynamic, collaborative, progressive, and high-performing team
  • Flexible work/life balance options
  • Opportunities to do challenging work
  • Opportunities to take on progressively greater accountabilities
  • Access to a variety of job opportunities across business


#LI-post

#TECHPJ

Job Skills
Agile Methodology, Ansible Tower, Group Problem Solving, IT System Administration, IT Systems Integration, Kubernetes, Linux, Organizational Leadership, Product Services, RedHat OpenShift Administration, Red Hat OS Administration, Software Development Life Cycle (SDLC), System Applications, System Integration Testing (SIT), Systems Software

Additional Job Details

Address:

RBC CENTRE, 155 WELLINGTON ST W:TORONTO

City:

Toronto

Country:

Canada

Work hours/week:

37.5

Employment Type:

Full time

Platform:

TECHNOLOGY AND OPERATIONS

Job Type:

Regular

Pay Type:

Salaried

Posted Date:

2026-07-31

Application Deadline:

2026-08-28
Note: Applications will be accepted until 11:59 PM on the day prior to the application deadline date above

About Royal Bank of Canada

Royal Bank of Canada Careers

Join the dynamic team at Royal Bank of Canada (RBC), a global leader in financial services and a company committed to excellence and innovation. At RBC, we offer a wide range of job opportunities that empower professionals to shape their career paths with leadership, diversity training, and continuous growth.

Work You’ll Do

At Royal Bank of Canada, we are not just hiring; we are building a culture of innovation and leadership. Our team members are at the forefront of the financial industry, driving transformation and delivering targeted solutions that meet the evolving needs of our clients and communities.

Explore Job Opportunities and Employment at RBC

Whether you are starting your career or looking to take it to the next level, RBC offers positions that challenge your skills and fuel your ambition. From entry-level positions to leadership roles, our job opportunities span across various functions and regions. Join us and be part of a team that values professional growth and diversity.

Internship and Professional Development

Kickstart your career with an internship at Royal Bank of Canada. Our internships provide invaluable hands-on experience, networking opportunities, and insights into the financial services industry. Interns at RBC gain the skills necessary to excel and are often considered for full-time positions within the company.

Benefits and Culture

At RBC, we prioritize the well-being and satisfaction of our employees. Our benefits package is designed to support our team members at every stage of their life and career. RBC’s culture is built on a foundation of respect, integrity, and responsibility, fostering an environment where everyone can thrive.

Career Growth and Innovation

We believe in nurturing the potential of our employees through continuous learning and career development programs. At RBC, you will find endless opportunities to grow professionally through on-the-job experiences, formal training programs, and leadership development initiatives. Our commitment to innovation means we are constantly seeking out new ideas and perspectives, making RBC a perfect place for those who aim to lead and innovate.

Diversity and Inclusion

Diversity is our strength. At Royal Bank of Canada, we are committed to building an inclusive workplace where every employee feels valued and respected. Our diversity training programs are designed to educate and inspire, creating a more inclusive and equitable workplace.

Join Our Team

Search open positions that match your skills and interests. We look for passionate, curious, creative, and solution-driven team players. Start your journey with RBC today and be part of a world-class team known for its commitment to client service, community involvement, and innovation.

Stay Connected

Keep up to date with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here at Royal Bank of Canada.

Job Alert Emails

Personalize your subscription to receive job alerts, latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities awaiting you at RBC. Explore the possibilities with Royal Bank of Canada, where your future is filled with potential and the path to success is paved with countless opportunities for professional and personal growth. Join us and shape not just your career but the future of the financial industry.
Learn more about Royal Bank of Canada
Size
86,007 employees
Market Cap
$130.3 billion
Industry
5 Year Trend
+8.7%
NASDAQ

Similar Jobs

More Jobs at Royal Bank of Canada

More Information Technology Jobs

Find similar Site Reliability Engineer (SRE), Cloud Operations jobs: