ABOUT THE ROLE & TEAMWe are seeking a
hands-on Site Reliability Engineer (SRE) with strong expertise across application support, Kubernetes environments, and CI/CD pipelines. This role is responsible for ensuring the reliability, performance, and observability of production systems through deep technical analysis and proactive engineering.
The successful candidate will lead incident and problem investigations, identify root causes, and drive permanent resolutions in close collaboration with Development, Product, and Operations teams. This is an engineering-focused role requiring strong troubleshooting skills, production support experience, and a commitment to continuous service improvement.
WHAT YOU WILL DO- Analyze production incidents using logs, metrics, and traces to identify impacted application code and execution paths.
- Diagnose system issues and determine whether the root cause is related to application code, configuration, Kubernetes, or infrastructure.
- Troubleshoot Kubernetes workloads, including runtime behavior, networking, health probes, and failure scenarios.
- Serve as the technical escalation point during critical incidents, providing clear and timely guidance.
- Lead root cause analysis (RCA) and drive permanent corrective actions to improve reliability.
- Enhance observability and alerting to improve issue detection and resolution.
- Partner with Development and Platform teams to resolve systemic issues and strengthen service reliability.
- Automate repetitive operational tasks and promote engineering best practices.
- Assess the impact of deployments on production environments through CI/CD pipeline expertise.
- Monitor system performance and reliability, identifying opportunities to improve resilience and reduce outages.
QualificationsABOUT YOUR SKILLS- 5+ years' experience in SRE, DevOps, or Production Engineering supporting high-availability systems.
- Strong expertise in root cause analysis (RCA) and permanent issue resolution.
- Hands-on troubleshooting of production environments using logs, metrics, and traces.
- Strong knowledge of .NET and/or Java applications in production.
- Hands-on experience with Kubernetes and containerized workloads.
- Experience with monitoring and observability tools for distributed systems.
- Familiarity with CI/CD pipelines and deployment tools (e.g., Azure DevOps, Jenkins, GitHub Actions).
- Scripting and automation skills using Python, Bash, or similar.
- Strong analytical, communication, and cross-functional collaboration skills.
- Bachelor's degree in Computer Science, Engineering, or related field, or equivalent experience.
WHAT WE OFFERWe value diversity, operating in 200 countries and spanning 60 languages and cultures. Our inclusive offices are comfortable and fun, with the flexibility to work from home. Join our team and step closer to your best life.
Flex Week: Hybrid (2 days from home and 3 days in the Montreal office.
🌎
Flex Location: Take up to 30 days a year to work from any location in the world.
🌿
Employee Wellbeing: We've got you covered with our Employee Assistance Program (EAP), for you and your dependents 24/7, 365 days/year. We also offer Champion Health a personalized platform that supports a range of well-being needs.
Professional Development: Level up your skills with our training platforms, including LinkedIn Learning!
Competitive Benefits: Competitive benefits that make sense with both your local market and employment status.