Site Reliability Engineering SRE - Lead

Compunnel

$150K — $180K *
Finance & Insurance
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of technical experience in SRE, DevOps, or Platform Engineering within regulated environments.
  • Proficiency in AWS and on-premises architecture design.
  • Expertise in observability tools including Prometheus and Grafana.
  • Hands-on experience with CI/CD processes and Infrastructure as Code.
  • Strong coding skills in Python or Go, with familiarity in automation tools like Ansible and Terraform.
  • Experience with defining SLIs, SLOs, and SLAs in a production environment.
  • Excellent communication skills to engage both technical and non-technical stakeholders.

Responsibilities

  • Own technical outcomes and enhance engineering quality standards across teams.
  • Mentor engineering staff and influence technical decisions through leadership.
  • Collaborate with various teams to integrate reliability throughout the software development lifecycle.
  • Define and monitor key performance indicators for reliability and efficiency.
  • Lead incident management processes, ensuring swift resolution and effective stakeholder communication.
  • Establish operational standards including runbooks and readiness protocols.
  • Drive continuous improvements in operational processes and incident responses.

Benefits

  • Flexible hybrid work arrangement with local consulting opportunities.
  • Opportunities for professional growth and skill enhancement.
  • Engagement in a high-stakes, mission-critical financial services environment.
  • Hands-on involvement with cutting-edge technology such as AI-assisted management.
  • Collaboration with multidisciplinary teams for integrated solutions.
Full Job Description
Job Summary
Seeking a highly technical Site Reliability Engineering (SRE) Lead to design and deliver highly available, scalable, and resilient platforms within a large-scale financial services environment. This role will deliver technical solutions that improve system stability, observability, and resilience while ensuring mission-critical systems meet stringent reliability, performance, and regulatory requirements. The SRE Lead will provide hands-on technical leadership across engineering, incident management, automation, observability, and resilience architecture while collaborating closely with product, platform, operations, security, development, and support teams.

Key Responsibilities
• Take ownership of technical outcomes, raise engineering quality standards, and share expertise across teams.
• Mentor peers, influence technical decisions through hands-on leadership, and build engineering capabilities through cross-team collaboration.
• Partner with product, platform, operations, security, development, and support teams to embed reliability throughout the software development lifecycle.
• Define and track KPIs for reliability, performance, and operational efficiency.
• Lead production incident management, including detection, triage, escalation, mitigation, and resolution.
• Manage major incident response for critical outages, ensuring rapid mitigation and clear stakeholder communication.
• Conduct blameless postmortems and ensure actionable follow-up items and systemic improvements are implemented.
• Establish runbooks, playbooks, and operational readiness standards for supported services.
• Drive continuous improvement in incident response processes, tooling, and operational readiness.
• Drive AI-assisted incident management, problem management, change management, and risk management across ITSM activities.
• Automate operational processes, including incident response, failover, scaling, and recovery.
• Implement self-healing mechanisms using auto-remediation, event-driven workflows, and AI/ML-assisted operations where applicable.
• Implement Infrastructure as Code (IaC) using tools such as Terraform and CloudFormation.
• Reduce manual operational effort through CI/CD pipelines, automated testing, and deployment strategies such as blue/green and canary releases.
• Implement observability frameworks covering metrics, logs, and traces using tools such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, CloudWatch, Datadog, or similar technologies.
• Use observability data to proactively identify and eliminate system risks and improve platform reliability.
• Define and implement SLIs, SLOs, SLAs, and error budgets across critical services.
• Apply resilient design patterns including multi-region failover, active-active architectures, circuit breakers, and bulkheads.
• Ensure platforms support high availability, fault tolerance, and disaster recovery requirements for mission-critical systems.
• Lead deep-dive investigations into platform failures, performance bottlenecks, and systemic issues across complex, multi-tiered architectures.
• Maintain effective communication with technical and non-technical stakeholders during incidents and operational activities.
• Apply a hands-on, strategic approach to reliability engineering, architectural clarity, and continuous improvement.

Required Qualifications
• Proven technical experience in SRE, DevOps, and/or Platform Engineering within large-scale, regulated environments.
• Strong experience with both on-premises and AWS cloud-native architecture and systems design.
• Strong implementation experience with observability tooling and frameworks for metrics, logs, and traces, including Prometheus, Grafana, OpenTelemetry, ELK/EFK, CloudWatch, Datadog, or similar technologies.
• In-depth experience with incident management and production operations.
• Hands-on experience with CI/CD pipelines, Git, automation, scripting, and Infrastructure as Code tools.
• Experience with Python, Go, Ansible, Terraform, or similar technologies.
• Demonstrable experience applying agentic AI engineering to observability, incident management, and automated recovery.
• Strong understanding of networking, security, and reliability engineering principles.
• Experience defining and implementing SLOs, SLIs, SLAs, and error budgets.
• Hands-on experience with AWS services including EKS/ECS, Lambda, API Gateway, DynamoDB, Aurora, and S3.
• Experience with database and data platforms including SQL Server, Sybase, PostgreSQL, and Aurora.
• Experience supporting and/or developing applications using Unix/Linux, Java, C#/.NET, and Python.
• Experience working in the financial services domain or another regulated environment.
• Strong communication skills with the ability to engage technical and non-technical stakeholders.
• Ability to remain calm, decisive, and effective during production incidents and high-pressure situations.
• Local consultant availability for the required hybrid work arrangement.

Preferred Qualifications
• Experience with financial services or other data-intensive, regulated industries.
• Familiarity with multi-region AWS architectures and highly available, mission-critical applications.
• Knowledge of chaos engineering practices and resilience testing.
• Exposure to AIOps and intelligent automation frameworks.

Similar Jobs

More Jobs at Compunnel

More Finance & Insurance Jobs

Find similar Site Reliability Engineering SRE - Lead jobs: