Relx Group

Site Reliability Engineering Lead

Relx Group • $118K — $219K *
US-Anywhere
+ 8 other locationsRemote
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in SRE, DevOps, or Infrastructure roles
  • Expert knowledge of Kubernetes and its architecture
  • Proficient in Terraform for infrastructure as code
  • Deep understanding of Azure Cloud services
  • Experience with CI/CD pipeline design and automation
  • Strong programming skills in Python, Bash, or PowerShell
  • Proven leadership in incident response and reliability improvements

Responsibilities

  • Manage and mentor a team of SREs, conducting performance reviews and career development
  • Oversee hiring, onboarding, and resource allocation for the team
  • Set team goals and prioritize project backlogs
  • Promote a blameless culture for post-incident reviews and collaboration
  • Lead initiatives to enhance reliability across services and infrastructure
  • Drive incident response and continuous service improvement efforts
  • Champion automation to enhance operational excellence

Benefits

  • Country-specific benefits tailored to employee well-being
  • Annual incentive bonus eligibility
  • Opportunities for personal development and career growth
  • Supportive team culture focused on collaboration and innovation
  • Access to modern cloud environments and cutting-edge technologies
Full Job Description
Are you passionate about building reliable, scalable platforms and helping engineering teams thrive?

Do you enjoy leading high-performing teams while driving automation, resilience, and operational excellence across modern cloud environments?

About Our Team

You will be joining the Core SRE Team in Business Services, a team that oversees all the applications and infrastructure in the biggest business unit in LexisNexis Risk Solutions. We build cloud environments, migrate on prem applications to the cloud, work with self-hosted and 3rd party solutions. The successful candidate is a self-starter who assesses the situation, collaboratively develops a solution, and takes the initiative to improve performance, cost, and reliability at each opportunity.

About the Role

This is a professional management level role. Individuals are required to provide line management to a small to medium sized team of engineers, including authority over performance management, pay, and recruitment. They will ensure that tasks and projects are prioritized appropriately and provide support to team members when tasks are blocked. They will support engineers in their personal development and ensure they are working within the SRE framework. They will lead the post-mortem reviews and the timely production of RCAs. They address issues with impact beyond their own team based on knowledge of related disciplines.

Responsibilities
  • Manage, mentor, and grow a team of SREs; conduct 1:1s, performance reviews, and career development planning
  • Own hiring, onboarding, and team capacity/resourcing decisions
  • Set team goals, prioritize backlog, and drive planning
  • Foster a blameless post-incident culture and cross-team collaboration with Dev, Security, and Product
  • Lead reliability initiatives across infrastructure and services
  • Drive incident response activities and continuous service improvement
  • Champion automation and operational excellence across the platform
  • Support the development of scalable, secure, and resilient cloud-native environments


Requirements
  • Expert knowledge of Kubernetes, including cluster architecture, upgrades, autoscaling, security hardening, and troubleshooting at scale
  • Expert experience with Terraform, including modular IaC design, state management, multi-environment provisioning, and policy-as-code
  • Deep knowledge of Azure Cloud, including compute, networking, identity (AAD), storage, and cost optimization
  • Experience designing and scaling CI/CD pipelines using GitHub Actions, release strategies, and rollback automation
  • Experience with observability platforms including Prometheus, Grafana, OpenTelemetry, and SLO/SLA/error-budget management
  • Strong automation skills, focused on eliminating toil through self-healing systems and infrastructure automation
  • Advanced proficiency in Python, Bash, and/or PowerShell for tooling and automation
  • Deep understanding of networking concepts including TCP/IP, DNS, load balancing, VPNs, and cloud-native networking
  • Experience in SRE, DevOps, or Infrastructure roles, including experience leading engineering teams
  • Proven track record leading incident response and driving reliability improvements


U.S. National Base Pay Range: $118,300 - $219,800. Geographic differentials may apply in some locations to better reflect local market rates.This job is eligible for an annual incentive bonus.
We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits. Click here to access benefits specific to your location.

About Relx Group

RELX Group is a global provider of information-based analytics and decision tools for professional and business customers. The company operates in four market segments: scientific, technical and medical; risk and business analytics; legal; and exhibitions. RELX's products and services include electronic databases, online information services, workflow tools, and print and digital books. The company was founded in 1993 and is headquartered in London, England.
Learn more about Relx Group
Size
33,500 employees
Market Cap
$53.1 billion
Industry
Net Income
$1.2 billion
Founded
2018
5 Year Trend
+1%
Revenue
$7.1 billion
NASDAQ

Similar Jobs

More Jobs at Relx Group

More Information Technology Jobs

Find similar Site Reliability Engineering Lead jobs: