OMERS Administration Corporation

Lead, Site Reliability Engineering (Application Support)

Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in Site Reliability Engineering, Platform Engineering, or similar roles
  • Hands-on experience with Microsoft Azure services like Azure Container Apps
  • Experience in production environment support and incident response
  • Familiarity with container technologies and cloud-native architectures
  • Knowledge of CI/CD and automation using GitHub Actions
  • Strong grasp of identity and access management concepts
  • Experience with monitoring tools like Datadog and Azure Monitor
  • Scripting experience with PowerShell or Python.

Responsibilities

  • Monitor and support applications across various environments.
  • Lead incident response and coordinate with cross-functional teams to resolve issues.
  • Assist with deployments, change management, and CI/CD pipelines.
  • Troubleshoot user access issues related to Azure AD and permissions.
  • Configure support for Azure platform components and services.
  • Analyze performance issues using observability tools like Log Analytics.
  • Guide onboarding of applications onto the DEV platform.

Benefits

  • Opportunity to work with cutting-edge technologies in a dynamic environment.
  • Access to training and development resources for continuous improvement.
  • Flexible hybrid work arrangement requiring office presence 4 days a week.
  • Participation in group benefits and retirement plans.
Full Job Description
Role Summary

The Lead, Site Reliability Engineering ensures monitoring and analysis is conducted to guarantee the ongoing stability of all systems. The Lead is also expected to implement and maintain the infrastructure and tools that are necessary to manage the software development process.

Reporting to the SRE, Developer Platform Engineering, the Lead, DEV Platform Support Engineer will play a critical role in supporting and enhancing our Azure-based developer platforms and internal applications. This position combines platform engineering, site reliability engineering, and production support responsibilities to ensure applications remain reliable, secure, and easy to operate.

This is an excellent opportunity for someone who is passionate about Site Reliability Engineering and enjoys improving the reliability, availability, performance, and operability of cloud-native platforms. The successful candidate will apply SRE practices such as incident response, observability, automation, runbook development, root-cause analysis, and continuous reliability improvement while partnering with cross-functional teams to onboard, deploy, monitor, and support applications across the organization.

You Will Be Responsible For
  • Monitor, troubleshoot, and support applications and developer platform services across DEV, UAT, and PROD environments.
  • Respond to incidents, lead triage activities, and work closely with SRE, Platform, Network, Security, and application teams to restore service and resolve issues.
  • Support deployments, release activities, change management, and CI/CD pipelines, including GitHub Actions workflows.
  • Check, troubleshoot, and resolve user and developer access issues, including Azure AD groups, SSO, application permissions, and firewall rules.
  • Configure and support platform components such as Azure Container Apps, App Registrations, Key Vault, DNS, certificates, networking, and shared cloud services.
  • Investigate performance, reliability, and availability issues using Datadog, Azure Monitor, Log Analytics, and related observability tools.
  • Support onboarding of new applications and teams to the DEV platform by helping with setup, access, deployment readiness, monitoring, and operational handover.
  • Develop and maintain runbooks, support procedures, knowledge articles, and operational documentation to improve support effectiveness and knowledge sharing.
  • Contribute to automation and continuous improvement initiatives that reduce manual effort, improve reliability, and strengthen operational processes.
  • Provide technical guidance to team members and stakeholders while promoting Site Reliability Engineering and platform support best practices.


Required Skills & Experience
  • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or Production Support.
  • Strong hands-on experience with Microsoft Azure services, including Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Storage Accounts, Azure SQL, API Management (APIM), and Azure Functions.
  • Experience supporting production environments, including incident response, troubleshooting, problem management, and operational support processes.
  • Experience with container technologies and cloud-native application architectures.
  • Hands-on experience with CI/CD pipelines and deployment automation using GitHub Actions or similar platforms.
  • Strong understanding of identity, networking, and access management concepts, including SSO, OAuth, application registrations, and security groups.
  • Experience with observability and monitoring platforms such as Datadog, Azure Monitor, and Log Analytics.
  • Understanding of cloud networking concepts, including DNS, certificates, firewalls, private endpoints, and network security controls.
  • Experience with scripting and automation using technologies such as PowerShell, Bash, Azure CLI, Python, or similar tools.
  • Strong knowledge of operating systems and cloud infrastructure concepts.
  • Proven ability to work effectively in cross-functional environments and collaborate with technical and business stakeholders.
  • Strong communication, problem-solving, and organizational skills.


Preferred Skills & Experience
  • Experience with container apps and Kubernetes container orchestration platforms.
  • Experience with different pipelines.
  • Experience supporting enterprise developer platforms or internal platform engineering teams.
  • Knowledge of Site Reliability Engineering principles, including SLOs, SLIs, error budgets, and reliability engineering practices.
  • Experience with enterprise API integrations and platform services.
  • Exposure to AI, LLM, or agent-based technology platforms.
  • Experience with Azure networking and security best practices in enterprise environments.
  • Azure, Network, DevOps, or cloud-related certifications.
  • Post-secondary education in Computer Science, Software Engineering, Information Technology, or a related discipline.


We believe that time together in the office is important for OMERS and Oxford, the strength of our employees, and the work we do for our pension members. In delivering on our pension promise, keeping us connected to our work and each other, our flexible hybrid work guideline requires teams to come in to the office 4 days per week.

This posting is for an existing vacancy.

The expected salary range for this position is $86,000.00 - $130,000.00 per year.

You may also be eligible to receive an annual Incentive Award pursuant to our Short-term Incentive plan and our Long-Term Incentive plan (if applicable), and to participate in our group benefits and retirement plans - details on these elements of compensation are included within OMERS & Oxford offer letters.

About OMERS Administration Corporation

OMERS Administration Corporation is a Canadian pension fund that manages investments for the Ontario Municipal Employees Retirement System (OMERS). OMERS is one of Canada's largest pension funds, with over 500,000 members and over CAD 100 billion in net assets. OMERS Administration Corporation manages a diversified portfolio of investments across various asset classes, including public equity, private equity, infrastructure, real estate, and fixed income. The company's mission is to provide secure and sustainable pensions to its members while generating returns that help fund their pensions. OMERS Administration Corporation is headquartered in Toronto, Canada.
Learn more about OMERS Administration Corporation
Size
2,700 employees
Industry

Similar Jobs

More Jobs at OMERS Administration Corporation

More Information Technology Jobs

Find similar Lead, Site Reliability Engineering (Application Support) jobs: