Job Summary
Seeking a highly technical Site Reliability Engineering (SRE) Lead to design and deliver highly available, scalable, and resilient platforms within a large-scale financial services environment. This role will deliver technical solutions that improve system stability, observability, and resilience while ensuring mission-critical systems meet stringent reliability, performance, and regulatory requirements. The SRE Lead will provide hands-on technical leadership across engineering, incident management, automation, observability, and resilience architecture while collaborating closely with product, platform, operations, security, development, and support teams.
Key Responsibilities
• Take ownership of technical outcomes, raise engineering quality standards, and share expertise across teams.
• Mentor peers, influence technical decisions through hands-on leadership, and build engineering capabilities through cross-team collaboration.
• Partner with product, platform, operations, security, development, and support teams to embed reliability throughout the software development lifecycle.
• Define and track KPIs for reliability, performance, and operational efficiency.
• Lead production incident management, including detection, triage, escalation, mitigation, and resolution.
• Manage major incident response for critical outages, ensuring rapid mitigation and clear stakeholder communication.
• Conduct blameless postmortems and ensure actionable follow-up items and systemic improvements are implemented.
• Establish runbooks, playbooks, and operational readiness standards for supported services.
• Drive continuous improvement in incident response processes, tooling, and operational readiness.
• Drive AI-assisted incident management, problem management, change management, and risk management across ITSM activities.
• Automate operational processes, including incident response, failover, scaling, and recovery.
• Implement self-healing mechanisms using auto-remediation, event-driven workflows, and AI/ML-assisted operations where applicable.
• Implement Infrastructure as Code (IaC) using tools such as Terraform and CloudFormation.
• Reduce manual operational effort through CI/CD pipelines, automated testing, and deployment strategies such as blue/green and canary releases.
• Implement observability frameworks covering metrics, logs, and traces using tools such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, CloudWatch, Datadog, or similar technologies.
• Use observability data to proactively identify and eliminate system risks and improve platform reliability.
• Define and implement SLIs, SLOs, SLAs, and error budgets across critical services.
• Apply resilient design patterns including multi-region failover, active-active architectures, circuit breakers, and bulkheads.
• Ensure platforms support high availability, fault tolerance, and disaster recovery requirements for mission-critical systems.
• Lead deep-dive investigations into platform failures, performance bottlenecks, and systemic issues across complex, multi-tiered architectures.
• Maintain effective communication with technical and non-technical stakeholders during incidents and operational activities.
• Apply a hands-on, strategic approach to reliability engineering, architectural clarity, and continuous improvement.
Required Qualifications
• Proven technical experience in SRE, DevOps, and/or Platform Engineering within large-scale, regulated environments.
• Strong experience with both on-premises and AWS cloud-native architecture and systems design.
• Strong implementation experience with observability tooling and frameworks for metrics, logs, and traces, including Prometheus, Grafana, OpenTelemetry, ELK/EFK, CloudWatch, Datadog, or similar technologies.
• In-depth experience with incident management and production operations.
• Hands-on experience with CI/CD pipelines, Git, automation, scripting, and Infrastructure as Code tools.
• Experience with Python, Go, Ansible, Terraform, or similar technologies.
• Demonstrable experience applying agentic AI engineering to observability, incident management, and automated recovery.
• Strong understanding of networking, security, and reliability engineering principles.
• Experience defining and implementing SLOs, SLIs, SLAs, and error budgets.
• Hands-on experience with AWS services including EKS/ECS, Lambda, API Gateway, DynamoDB, Aurora, and S3.
• Experience with database and data platforms including SQL Server, Sybase, PostgreSQL, and Aurora.
• Experience supporting and/or developing applications using Unix/Linux, Java, C#/.NET, and Python.
• Experience working in the financial services domain or another regulated environment.
• Strong communication skills with the ability to engage technical and non-technical stakeholders.
• Ability to remain calm, decisive, and effective during production incidents and high-pressure situations.
• Local consultant availability for the required hybrid work arrangement.
Preferred Qualifications
• Experience with financial services or other data-intensive, regulated industries.
• Familiarity with multi-region AWS architectures and highly available, mission-critical applications.
• Knowledge of chaos engineering practices and resilience testing.
• Exposure to AIOps and intelligent automation frameworks.