Job Description
This is a hands-on engineering role with strong leadership influence, focused on reliability and platform. You will lead through technical credibility and influence setting standards, shaping direction, and elevating the reliability capability across engineering while remaining deeply hands-on. As TII's Lead Infrastructure & Site Reliability Engineer, you will own the run-time reliability, observability, cloud security, and platform engineering that keep our all-Azure environment secure, resilient, and always-on supporting the platform and business.
This is a build-and-enhance role: you will mature our observability across metrics, logs, and traces; establish disciplined incident command and blameless post-incident practices; codify infrastructure with Bicep and Terraform; harden the security posture of our Azure platform; and build the self-service platform capabilities that let engineering teams move fast safely. You will own reliability, security, and platform hands-on while partnering with engineering architecture, engineering leadership, and security stakeholders. With the AVP, Infrastructure, co-create the future-state vision and roadmap for Cloud, Observability, ITSM, and related domains making it a shared, measurable engineering discipline in support of TII's growth target and expansion into new distribution channels.
What you will do:
Infrastructure Strategy
- Own and evolve TII's Azure infrastructure and reliability roadmap, aligned to the Azure Well-Architected Framework.
- With the AVP, Infrastructure to co-create the future-state vision and roadmap for Cloud, Observability, ITSM, and related domains.
- Define standards for compute, network, and platform services that scale with business growth and new channels.
- Drive cloud cost optimization (FinOps) balancing performance, resilience, and spend.
- Partner with engineering architecture to ensure infrastructure enables design-time resilience and delivery velocity.
Site Reliability Engineering
- Own run-time reliability across availability, performance, scalability, and capacity for TII's platform.
- Mature and expand the SLO practice defining SLIs, refining the 99.9% (and higher, where warranted) SLOs, and operating error budgets to balance reliability and delivery speed.
- Lead capacity planning and performance engineering to support the platform's growth.
- Drive operational readiness reviews for new services and major releases.
Observability
- Own and mature the observability platform across the three pillars metrics, logs, and traces enhancing Grafana/Prometheus and Azure Application Insights.
- Implement distributed tracing across the GraphQL/REST services to accelerate diagnosis and reduce time-to-detect and time-to-resolve.
- Establish meaningful alerting and telemetry that reduce noise and surface real signals.
- Build reliability dashboards that give teams and leadership clear visibility into service health.
Platform Engineering
- Own the internal developer platform and self-service infrastructure capabilities that enable engineering teams to provision and operate safely.
- Define golden paths / paved-road templates that make the reliable, secure, and compliant way the easy way.
- Establish and champion Infrastructure as Code standards using Bicep and Terraform.
- Improve developer experience and engineering enablement through automation and reusable platform services.
Operational Excellence & Automation
- Drive automation across provisioning, configuration, deployment, and remediation to eliminate toil.
- Establish operational runbooks, self-healing patterns, and proactive reliability practices.
- Continuously improve deployment safety and rollback capability in partnership with CI/CD owners.
Cloud Security & Compliance
- Own the engineering and operational security of the Azure cloud platform including identity and access management, network security, configuration hardening, and secrets/key management.
- Manage and improve the cloud security posture (e.g., Microsoft Defender for Cloud, Azure Policy), including continuous vulnerability management and remediation.
- Implement security monitoring and alerting as part of the observability platform to detect and respond to threats.
- Embed DevSecOps and secure-by-design practices into platform and IaC workflows, enforcing guardrails and policy-as-code within golden paths and self-service tooling.
- Partner with the Security function on policy, governance, and compliance in a regulated insurance (PII) environment.
Incident & Problem Management
- Own the major-incident process and incident command, matured on the on-call platform (Better Stack).
- Lead blameless post-incident reviews and drive systemic problem management to prevent recurrence.
- Improve on-call health, escalation paths, and mean-time-to-detect / mean-time-to-resolve.
- Maintain and enhance the mature Business Continuity and Disaster Recovery capabilities, including RTO/RPO targets and periodic testing.
Leadership
- Will co-lead a small team of infrastructure, cloud, & system engineers.
- Uplift the reliability and platform capability across engineering raising standards and building a reliability culture.
- Mentor engineers, demonstrating the leadership behaviors that support growth into a formal infrastructure leadership role.
- Establish standards, documentation, and ways of working that scale across teams.
- Other duties as assigned
What YOU will bring to C&F:
- Excellent problem-solving and analytical skills with attention to detail
- Strong analytical and problem-solving skills.
- Excellent verbal and written communication skills, with the ability to explain technical and functional issues clearly to both technical and non-technical stakeholders.
- Demonstrated leadership or mentoring of distributed/offshore teams (formal people-management experience a plus).
- Excellent collaboration and influencing skills
- Outcome & Metrics Orientation
- Self-starter
Requirements:
- Bachelor's degree in Computer Science, or related field or equivalent experience.
- 8+ years in infrastructure, site reliability, platform, or DevOps engineering, with recent hands-on delivery.
- Deep, hands-on expertise operating production workloads on Microsoft Azure.
- Proven experience with Infrastructure as Code using Bicep and/or Terraform.
- Hands-on experience with observability tooling across metrics, logs, and traces (e.g., Grafana, Prometheus, Azure Application Insights).
- Proven experience defining and operating SLIs/SLOs and error budgets.
- Hands-on experience securing Azure cloud environments identity/access, network security, posture management (e.g., Microsoft Defender for Cloud, Azure Policy), and vulnerability management.
- Experience scaling Microsoft Fabric and Microsoft Purview.
- Experience owning incident management, on-call, and blameless post-incident reviews (e.g., Better Stack, PagerDuty, or comparable).
- Strong scripting/automation ability (e.g., PowerShell, Python, Bash) and CI/CD experience (Azure DevOps preferred).
- Experience supporting business continuity and disaster recovery with defined RTO/RPO.
- Experience building internal developer platforms, golden paths, and self-service infrastructure.
- Cloud cost optimization / FinOps experience.
- Exposure to AI-assisted operations (AIOps) and modern reliability automation.
- Microsoft Azure certifications (e.g., Azure Solutions Architect Expert, Azure DevOps Engineer Expert, Azure Security Engineer Associate) preferred.
- Experience in insurance, travel, fintech, or SaaS (regulated environments) preferred.
- Prior leadership or mentoring of distributed/offshore teams desired. Formal people-leadership experience a plus
What C&F will bring to you
What C&F will bring to YOU:
- Competitive compensation package
- Generous 401K employer match
- Employee Stock Purchase plan with employer matching
- Generous Paid Time Off
- Excellent benefits that go beyond health, dental & vision. Our programs are focused on your whole familys wellness including your physical, mental and financial wellbeing
- A core C&F tenant is owning your career development so we provide a wealth of ways for you to keep learning, including tuition reimbursement, industry related certifications and professional training to keep you progressing on your chosen path
- A dynamic, ambitious, fun and exciting work environment
- We believe you do well by doing good and want to encourage a spirit of social and community responsibility, matching donation program, volunteer opportunities, and an employee driven corporate giving program that lets you participate and support your community
#LI-MU
#LI-Remote