Lead Infrastructure & Site Reliability Engineering

Crum & Forster

$105K — $198K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science or related field, or equivalent experience.
  • 8+ years in infrastructure, site reliability, platform, or DevOps engineering, focusing on hands-on delivery.
  • Deep expertise with production workloads operating on Microsoft Azure.
  • Proficiency in Infrastructure as Code using Bicep and/or Terraform.
  • Demonstrated experience with observability tools like Grafana, Prometheus, and Azure Application Insights.
  • Hands-on security management for Azure environments and vulnerability management processes.
  • Experience supporting business continuity with defined RTO/RPO.

Responsibilities

  • Own and evolve TII's Azure infrastructure and reliability roadmap.
  • Co-create future-state vision for Cloud, Observability, and ITSM with AVP, Infrastructure.
  • Define standards for scalable compute, network, and platform services.
  • Drive cloud cost optimization while balancing performance and spend.
  • Own run-time reliability for platform availability and performance.
  • Mature the Service Level Objective (SLO) practice and manage error budgets.
  • Establish golden paths and self-service infrastructure capabilities for engineering teams.

Benefits

  • Competitive compensation package.
  • Generous 401K employer match.
  • Employee Stock Purchase plan with employer matching.
  • Generous Paid Time Off.
  • Comprehensive health, dental, and vision benefits focused on wellness.
  • Tuition reimbursement and support for professional development.
  • Dynamic and exciting work environment with community engagement opportunities.
Full Job Description
Job Description

This is a hands-on engineering role with strong leadership influence, focused on reliability and platform. You will lead through technical credibility and influence-setting standards, shaping direction, and elevating the reliability capability across engineering-while remaining deeply hands-on. As TII's Lead Infrastructure & Site Reliability Engineer, you will own the run-time reliability, observability, cloud security, and platform engineering that keep our all-Azure environment secure, resilient, and always-on-supporting the platform and business.

This is a build-and-enhance role: you will mature our observability across metrics, logs, and traces; establish disciplined incident command and blameless post-incident practices; codify infrastructure with Bicep and Terraform; harden the security posture of our Azure platform; and build the self-service platform capabilities that let engineering teams move fast safely. You will own reliability, security, and platform hands-on while partnering with engineering architecture, engineering leadership, and security stakeholders. With the AVP, Infrastructure, co-create the future-state vision and roadmap for Cloud, Observability, ITSM, and related domains making it a shared, measurable engineering discipline in support of TII's growth target and expansion into new distribution channels.

What you will do:

Infrastructure Strategy
  • Own and evolve TII's Azure infrastructure and reliability roadmap, aligned to the Azure Well-Architected Framework.
  • With the AVP, Infrastructure to co-create the future-state vision and roadmap for Cloud, Observability, ITSM, and related domains.
  • Define standards for compute, network, and platform services that scale with business growth and new channels.
  • Drive cloud cost optimization (FinOps)-balancing performance, resilience, and spend.
  • Partner with engineering architecture to ensure infrastructure enables design-time resilience and delivery velocity.

Site Reliability Engineering
  • Own run-time reliability across availability, performance, scalability, and capacity for TII's platform.
  • Mature and expand the SLO practice-defining SLIs, refining the 99.9% (and higher, where warranted) SLOs, and operating error budgets to balance reliability and delivery speed.
  • Lead capacity planning and performance engineering to support the platform's growth.
  • Drive operational readiness reviews for new services and major releases.

Observability
  • Own and mature the observability platform across the three pillars-metrics, logs, and traces-enhancing Grafana/Prometheus and Azure Application Insights.
  • Implement distributed tracing across the GraphQL/REST services to accelerate diagnosis and reduce time-to-detect and time-to-resolve.
  • Establish meaningful alerting and telemetry that reduce noise and surface real signals.
  • Build reliability dashboards that give teams and leadership clear visibility into service health.

Platform Engineering
  • Own the internal developer platform and self-service infrastructure capabilities that enable engineering teams to provision and operate safely.
  • Define golden paths / paved-road templates that make the reliable, secure, and compliant way the easy way.
  • Establish and champion Infrastructure as Code standards using Bicep and Terraform.
  • Improve developer experience and engineering enablement through automation and reusable platform services.

Operational Excellence & Automation
  • Drive automation across provisioning, configuration, deployment, and remediation to eliminate toil.
  • Establish operational runbooks, self-healing patterns, and proactive reliability practices.
  • Continuously improve deployment safety and rollback capability in partnership with CI/CD owners.

Cloud Security & Compliance
  • Own the engineering and operational security of the Azure cloud platform-including identity and access management, network security, configuration hardening, and secrets/key management.
  • Manage and improve the cloud security posture (e.g., Microsoft Defender for Cloud, Azure Policy), including continuous vulnerability management and remediation.
  • Implement security monitoring and alerting as part of the observability platform to detect and respond to threats.
  • Embed DevSecOps and secure-by-design practices into platform and IaC workflows, enforcing guardrails and policy-as-code within golden paths and self-service tooling.
  • Partner with the Security function on policy, governance, and compliance in a regulated insurance (PII) environment.

Incident & Problem Management
  • Own the major-incident process and incident command, matured on the on-call platform (Better Stack).
  • Lead blameless post-incident reviews and drive systemic problem management to prevent recurrence.
  • Improve on-call health, escalation paths, and mean-time-to-detect / mean-time-to-resolve.
  • Maintain and enhance the mature Business Continuity and Disaster Recovery capabilities, including RTO/RPO targets and periodic testing.

Leadership
  • Will co-lead a small team of infrastructure, cloud, & system engineers.
  • Uplift the reliability and platform capability across engineering-raising standards and building a reliability culture.
  • Mentor engineers, demonstrating the leadership behaviors that support growth into a formal infrastructure leadership role.
  • Establish standards, documentation, and ways of working that scale across teams.
  • Other duties as assigned


What YOU will bring to C&F:
  • Excellent problem-solving and analytical skills with attention to detail
  • Strong analytical and problem-solving skills.
  • Excellent verbal and written communication skills, with the ability to explain technical and functional issues clearly to both technical and non-technical stakeholders.
  • Demonstrated leadership or mentoring of distributed/offshore teams (formal people-management experience a plus).
  • Excellent collaboration and influencing skills
  • Outcome & Metrics Orientation
  • Self-starter


Requirements:
  • Bachelor's degree in Computer Science, or related field-or equivalent experience.
  • 8+ years in infrastructure, site reliability, platform, or DevOps engineering, with recent hands-on delivery.
  • Deep, hands-on expertise operating production workloads on Microsoft Azure.
  • Proven experience with Infrastructure as Code using Bicep and/or Terraform.
  • Hands-on experience with observability tooling across metrics, logs, and traces (e.g., Grafana, Prometheus, Azure Application Insights).
  • Proven experience defining and operating SLIs/SLOs and error budgets.
  • Hands-on experience securing Azure cloud environments-identity/access, network security, posture management (e.g., Microsoft Defender for Cloud, Azure Policy), and vulnerability management.
  • Experience scaling Microsoft Fabric and Microsoft Purview.
  • Experience owning incident management, on-call, and blameless post-incident reviews (e.g., Better Stack, PagerDuty, or comparable).
  • Strong scripting/automation ability (e.g., PowerShell, Python, Bash) and CI/CD experience (Azure DevOps preferred).
  • Experience supporting business continuity and disaster recovery with defined RTO/RPO.
  • Experience building internal developer platforms, golden paths, and self-service infrastructure.
  • Cloud cost optimization / FinOps experience.
  • Exposure to AI-assisted operations (AIOps) and modern reliability automation.
  • Microsoft Azure certifications (e.g., Azure Solutions Architect Expert, Azure DevOps Engineer Expert, Azure Security Engineer Associate) preferred.
  • Experience in insurance, travel, fintech, or SaaS (regulated environments) preferred.
  • Prior leadership or mentoring of distributed/offshore teams desired. Formal people-leadership experience a plus


What C&F will bring to you

What C&F will bring to YOU:
  • Competitive compensation package
  • Generous 401K employer match
  • Employee Stock Purchase plan with employer matching
  • Generous Paid Time Off
  • Excellent benefits that go beyond health, dental & vision. Our programs are focused on your whole family's wellness including your physical, mental and financial wellbeing
  • A core C&F tenant is owning your career development so we provide a wealth of ways for you to keep learning, including tuition reimbursement, industry related certifications and professional training to keep you progressing on your chosen path
  • A dynamic, ambitious, fun and exciting work environment
  • We believe you do well by doing good and want to encourage a spirit of social and community responsibility, matching donation program, volunteer opportunities, and an employee driven corporate giving program that lets you participate and support your community


#LI-MU

#LI-Remote

Similar Jobs

More Jobs at Crum & Forster

More Information Technology Jobs

Find similar Lead Infrastructure & Site Reliability Engineering jobs: