Intact Financial Corporation

SRE specialist

Intact Financial Corporation$109K — $134K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years experience in SRE/Platform/Infrastructure/Software Engineering with large-scale systems
  • Strong skills in observability tools like OpenTelemetry, Dynatrace, and Elastic
  • Expertise in reliability engineering principles including SLIs, SLOs, and error budgets
  • Familiarity with CI/CD deployment patterns and Kubernetes
  • Solid programming skills in Go, Python, or TypeScript, with IaC experience (Terraform)
  • Experience with chaos engineering and disaster recovery exercises
  • Bilingual in French and English, with strong communication skills

Responsibilities

  • Lead high-severity incident investigations with multidisciplinary teams
  • Identify systemic risks and enhance resilience measures
  • Conduct blameless post-mortems and provide coaching
  • Implement observability solutions for metrics and tracing
  • Create auto-healing tooling and progressive delivery mechanisms
  • Define user-centric SLIs/SLOs and report on reliability performance
  • Upskill teams through standardized playbooks and training

Benefits

  • Flexible work arrangements with a hybrid work model
  • Opportunity to purchase up to 5 extra days off each year
  • Comprehensive benefits for physical and mental wellbeing including telemedicine
  • Employee share plan with matching contributions
  • Pension options providing long-term security and guaranteed income for life
Full Job Description

Pay at Intact is about much more than just salary.

  • Flexible work arrangements and a hybrid work model

  • Possibility to purchase up to 5 extra days off per year

  • Multiple benefits offered to support physical and mental wellbeing, including telemedicine, Wellness account and much more

  • Share plan & other savings: up to 12% of salary or even more (ask how you could earn guaranteed income for life)

Salary range (but not limited to):

109,900 - 134,300

Annual bonus target, based on the base salary, with a potential payout of up to double the target (subject to personal and company performance):

15%

As part of our commitment to Win As A Team, we share our success with employees through our annual bonus plan and Employee Share Purchase Plan (ESPP) 6 with Intact matching 50% of your net shares.

Our pension offerings provide flexibility and long-term security for our employees beyond their careers. We are one of the few companies offering the opportunity to receive guaranteed income for life via our defined benefit pension plan.

Salary for the candidate will be determined taking into consideration a number of factors including: experience, skills, qualifications, anticipated contribution to role, internal equity, etc. The salary range presented above is based on a 35-hour workweek and would represent a majority of different candidate profiles. However, we encourage candidates who may fall outside of this range to apply as well.


About the role

We are seeking a hands-on Site Reliability Engineer within the Intelligent Operations Department 27s SRE & Resiliency team. This role operates across Azure, AWS, GCP, and onprem environments, embedded in the broader enterprise resiliency and production reliability strategy. The SRE will function as part of a special investigations unit that empowers and enables Applicative Support, Infrastructure Support, and the Incident Management team 6 coaching, guiding, and leading investigations into active incidents and proactive reliability improvements. Core responsibilities include deep investigations, advanced observability (OpenTelemetry, Dynatrace, Elastic), auto-healing tooling, SLI/SLO stewardship, and business-aligned reliability reporting.

What you9ll do here:

Incidents & Investigations

  • Lead highseverity investigations and RCA with App/Infra/Incident teams.

  • Proactively find systemic risks and resilience gaps; drive durable fixes.

  • Run blameless postmortems and coach teams.

Observability (OTel, Dynatrace, Elastic)

  • Implement endtoend traces/metrics/logs with consistent semantics.

  • Build insights and anomaly detection; create topologyaware health models.

  • Integrate synthetics, contract tests, and distributed tracing.

AutoHealing & Reliability Tooling

  • Build policydriven remediation (circuit breakers, throttling, retries).

  • Enable progressive delivery (blue/green, canary) with safe rollbacks.

  • Provide resilience tooling: validation, safeguards, chaos, DR, runbooks.

SLI/SLOs & Reporting

  • Define usercentric SLIs/SLOs; enforce error budget policies.

  • Publish reliability reports and scorecards; drive continuous improvement.

Coaching & Leadership

  • Upskill support/incident teams; standardize playbooks and training.

  • Promote automationfirst, datadriven, resilience culture.

Cloud & Platform Reliability

  • Operate across Azure/AWS/GCP/onprem; GLB, DNS, TLS, CDN, failover.

  • Improve K8s/mesh (AKS/EKS/GKE, Istio/Linkerd) and data/streaming resilience.

AI for Reliability

  • Use AI for causal detection/anomalies to cut MTTR.

  • Develop reliability copilots; monitor AI systems for reliability and cost.

What you bring to the table:

  • 8+ years of experience in SRE/Platform/Infrastructure/Software Engineering operating large-scale production systems across multi-cloud and onprem.

  • Strong proficiency in:

    • Observability: OpenTelemetry instrumentation and standards; Dynatrace (Davis AI, SmartScape, service-level analysis, baselining); Elastic/ELK (Beats/Agent, ingest pipelines, ILM, Kibana).

    • Reliability engineering: SLIs/SLOs/SLAs, error budgets, alert strategy, capacity modeling, graceful degradation, circuit breaking, retries/backoff.

    • CI/CD and deployment patterns: blue/green, canary, progressive delivery, automated rollback, pipeline safeguards.

    • Kubernetes and service meshes; platform-level resilience and operability.

    • Data and event systems: replication, snapshots/PITR, CDC, streaming (Kafka, RabbitMQ, Pub/Sub) with DLQs/reprocessing; dependency risk management.

    • Networking and traffic: DNS, load balancers, CDN/edge, TLS/mTLS; fundamentals of BGP and global traffic management.

  • Solid software engineering skills in at least one of: Go, Python, or TypeScript; experience with IaC (Terraform), GitOps (Argo CD/Flux), and policy-as-code.

  • Experience running chaos engineering, game days, and DR exercises; ability to design safe experiments and embed learnings into production hardening.

  • Excellent communication (written, visual, verbal); adept at coaching, leading investigations, and presenting to mixed technical/business audiences.

  • Bilingual (French and English): Need to interact on a regular basis with an English-speaking clientele and colleagues across the country.

  • No Canadian work experiencerequiredhowever must be eligible to work in Canada

#LI-Hybrid

Il s9agit d9un nouveau rle au sein de notre quipe en plein croissance | This role is a new member of our growing team.

About Intact Financial Corporation

Intact Financial Corporation is a Canadian insurance company that provides property and casualty insurance to individuals and businesses. The company operates in Canada and the United States and offers a range of insurance products, including auto, home, and commercial insurance. Intact Financial Corporation was founded in 1809 and is headquartered in Toronto, Canada.
Learn more about Intact Financial Corporation
Size
16,000 employees
Industry

Similar Jobs

More Jobs at Intact Financial Corporation

More Information Technology Jobs

Find similar SRE specialist jobs: