Job DescriptionThis is a hands-on engineering role with strong leadership influence, focused on reliability and platform. You will lead through technical credibility and influence-setting standards, shaping direction, and elevating the reliability capability across engineering-while remaining deeply hands-on. As TII's Lead Infrastructure & Site Reliability Engineer, you will own the run-time reliability, observability, cloud security, and platform engineering that keep our all-Azure environment secure, resilient, and always-on-supporting the platform and business.
This is a build-and-enhance role: you will mature our observability across metrics, logs, and traces; establish disciplined incident command and blameless post-incident practices; codify infrastructure with Bicep and Terraform; harden the security posture of our Azure platform; and build the self-service platform capabilities that let engineering teams move fast safely. You will own reliability, security, and platform hands-on while partnering with engineering architecture, engineering leadership, and security stakeholders. With the AVP, Infrastructure, co-create the future-state vision and roadmap for Cloud, Observability, ITSM, and related domains making it a shared, measurable engineering discipline in support of TII's growth target and expansion into new distribution channels.
What you will do:Infrastructure Strategy- Own and evolve TII's Azure infrastructure and reliability roadmap, aligned to the Azure Well-Architected Framework.
- With the AVP, Infrastructure to co-create the future-state vision and roadmap for Cloud, Observability, ITSM, and related domains.
- Define standards for compute, network, and platform services that scale with business growth and new channels.
- Drive cloud cost optimization (FinOps)-balancing performance, resilience, and spend.
- Partner with engineering architecture to ensure infrastructure enables design-time resilience and delivery velocity.
Site Reliability Engineering- Own run-time reliability across availability, performance, scalability, and capacity for TII's platform.
- Mature and expand the SLO practice-defining SLIs, refining the 99.9% (and higher, where warranted) SLOs, and operating error budgets to balance reliability and delivery speed.
- Lead capacity planning and performance engineering to support the platform's growth.
- Drive operational readiness reviews for new services and major releases.
Observability- Own and mature the observability platform across the three pillars-metrics, logs, and traces-enhancing Grafana/Prometheus and Azure Application Insights.
- Implement distributed tracing across the GraphQL/REST services to accelerate diagnosis and reduce time-to-detect and time-to-resolve.
- Establish meaningful alerting and telemetry that reduce noise and surface real signals.
- Build reliability dashboards that give teams and leadership clear visibility into service health.
Platform Engineering- Own the internal developer platform and self-service infrastructure capabilities that enable engineering teams to provision and operate safely.
- Define golden paths / paved-road templates that make the reliable, secure, and compliant way the easy way.
- Establish and champion Infrastructure as Code standards using Bicep and Terraform.
- Improve developer experience and engineering enablement through automation and reusable platform services.
Operational Excellence & Automation- Drive automation across provisioning, configuration, deployment, and remediation to eliminate toil.
- Establish operational runbooks, self-healing patterns, and proactive reliability practices.
- Continuously improve deployment safety and rollback capability in partnership with CI/CD owners.
Cloud Security & Compliance- Own the engineering and operational security of the Azure cloud platform-including identity and access management, network security, configuration hardening, and secrets/key management.
- Manage and improve the cloud security posture (e.g., Microsoft Defender for Cloud, Azure Policy), including continuous vulnerability management and remediation.
- Implement security monitoring and alerting as part of the observability platform to detect and respond to threats.
- Embed DevSecOps and secure-by-design practices into platform and IaC workflows, enforcing guardrails and policy-as-code within golden paths and self-service tooling.
- Partner with the Security function on policy, governance, and compliance in a regulated insurance (PII) environment.
Incident & Problem Management- Own the major-incident process and incident command, matured on the on-call platform (Better Stack).
- Lead blameless post-incident reviews and drive systemic problem management to prevent recurrence.
- Improve on-call health, escalation paths, and mean-time-to-detect / mean-time-to-resolve.
- Maintain and enhance the mature Business Continuity and Disaster Recovery capabilities, including RTO/RPO targets and periodic testing.
Leadership- Will co-lead a small team of infrastructure, cloud, & system engineers.
- Uplift the reliability and platform capability across engineering-raising standards and building a reliability culture.
- Mentor engineers, demonstrating the leadership behaviors that support growth into a formal infrastructure leadership role.
- Establish standards, documentation, and ways of working that scale across teams.
- Other duties as assigned
What YOU will bring to C&F:- Excellent problem-solving and analytical skills with attention to detail
- Strong analytical and problem-solving skills.
- Excellent verbal and written communication skills, with the ability to explain technical and functional issues clearly to both technical and non-technical stakeholders.
- Demonstrated leadership or mentoring of distributed/offshore teams (formal people-management experience a plus).
- Excellent collaboration and influencing skills
- Outcome & Metrics Orientation
- Self-starter
Requirements:- Bachelor's degree in Computer Science, or related field-or equivalent experience.
- 8+ years in infrastructure, site reliability, platform, or DevOps engineering, with recent hands-on delivery.
- Deep, hands-on expertise operating production workloads on Microsoft Azure.
- Proven experience with Infrastructure as Code using Bicep and/or Terraform.
- Hands-on experience with observability tooling across metrics, logs, and traces (e.g., Grafana, Prometheus, Azure Application Insights).
- Proven experience defining and operating SLIs/SLOs and error budgets.
- Hands-on experience securing Azure cloud environments-identity/access, network security, posture management (e.g., Microsoft Defender for Cloud, Azure Policy), and vulnerability management.
- Experience scaling Microsoft Fabric and Microsoft Purview.
- Experience owning incident management, on-call, and blameless post-incident reviews (e.g., Better Stack, PagerDuty, or comparable).
- Strong scripting/automation ability (e.g., PowerShell, Python, Bash) and CI/CD experience (Azure DevOps preferred).
- Experience supporting business continuity and disaster recovery with defined RTO/RPO.
- Experience building internal developer platforms, golden paths, and self-service infrastructure.
- Cloud cost optimization / FinOps experience.
- Exposure to AI-assisted operations (AIOps) and modern reliability automation.
- Microsoft Azure certifications (e.g., Azure Solutions Architect Expert, Azure DevOps Engineer Expert, Azure Security Engineer Associate) preferred.
- Experience in insurance, travel, fintech, or SaaS (regulated environments) preferred.
- Prior leadership or mentoring of distributed/offshore teams desired. Formal people-leadership experience a plus
What C&F will bring to youWhat C&F will bring to YOU:- Competitive compensation package
- Generous 401K employer match
- Employee Stock Purchase plan with employer matching
- Generous Paid Time Off
- Excellent benefits that go beyond health, dental & vision. Our programs are focused on your whole family's wellness including your physical, mental and financial wellbeing
- A core C&F tenant is owning your career development so we provide a wealth of ways for you to keep learning, including tuition reimbursement, industry related certifications and professional training to keep you progressing on your chosen path
- A dynamic, ambitious, fun and exciting work environment
- We believe you do well by doing good and want to encourage a spirit of social and community responsibility, matching donation program, volunteer opportunities, and an employee driven corporate giving program that lets you participate and support your community
#LI-MU
#LI-Remote