Relx Group

Senior Site Reliability Engineer I

Relx Group • $95K — $158K *
US-AnywhereRemote in United States
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Advanced expertise in Terraform modules, state management, and lifecycle controls.
  • Hands-on experience with AWS operations in a multi-account and multi-region setup.
  • Proficient in GitHub Actions for CI/CD and deployment pipelines.
  • Knowledgeable in ECS Fargate, Docker, and container orchestration.
  • Strong understanding of AWS networking and cloud security best practices.
  • Experienced in incident response and system observability methodologies.
  • Proficient in Linux, automation scripting, and AWS CLI using Bash/Python scripts.
  • Hands-on experience in deploying and integrating AI tools in production environments.
  • Skilled in supporting developer enablement across engineering teams.

Responsibilities

  • Create monitoring queries and establish service level baselines.
  • Support senior engineers during incidents and troubleshooting.
  • Contribute to post-mortems and root cause analyses (RCAs).
  • Participate in disaster recovery testing initiatives.
  • Implement automation and execute code within production environments.
  • Contribute to the SRE knowledge base documentation.
  • Support deployment, monitoring, and reliability of AI-integrated services.
  • Assist in creating infrastructure topology drawings and deployment workflows.
  • Test system availability, reliability, and recovery in non-production environments.

Benefits

  • Country-specific benefits tailored to promote well-being and job satisfaction.
  • Potential eligibility for an annual incentive bonus.
Full Job Description
Senior Site Reliability Engineer

Are you passionate about building resilient, scalable systems that power mission-critical applications?
Do you thrive on automating operations, improving reliability, and ensuring exceptional system performance?

About the role:

As a Senior Site Reliability Engineer (SRE), you will play a key role in ensuring the reliability, scalability, and performance of our critical platforms and services. You will lead complex reliability initiatives, drive automation efforts to reduce operational toil, and help build resilient systems that deliver exceptional customer experiences.

You will leverage your expertise in observability, incident response, and distributed systems to proactively identify and resolve reliability challenges. Working closely with engineering teams, you will design and implement solutions that improve service availability, streamline operations, and enhance system recovery capabilities.

You will hold a high bar on code quality, flag risks and blockers early, and work alongside host-function stakeholders to make sure what you build fits real workflows, not assumed ones. You will also support handover and capability-building so the solution is owned and operable after the squad moves on.

Key Responsibilities:

  • Creating monitoring queries and establishes service level baselines.
  • Supporting senior engineers during incidents.
  • Making contributions during post-mortems and RCAs.
  • Participating in disaster recovery tests.
  • Implementing automation and executes code in production environments.
  • Contributing to SRE knowledge documentation.
  • Supporting the deployment, monitoring, and reliability of services integrating AI tools.
  • Supporting architecture and senior engineers in the creation of infrastructure topology drawings and deployment workflows.
  • Carrying out the testing of availability, reliability, and recoverability in non-production environments.


Requirements:

  • Advanced Terraform: Expertise in modules, providers, state management, lifecycle controls, drift detection, safe refactoring, and remote state (S3, locking, cross-stack dependencies).
  • AWS Operations: Hands-on experience managing production, multi-account, multi-region AWS environments across ECS, RDS, ALB, VPC, IAM, Route53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch.
  • GitHub Actions CI/CD: Experience building and troubleshooting reusable workflows, OIDC authentication, approval gates, runners, Terraform deployments, application deployments, and migration pipelines.
  • ECS Fargate & Containers: Knowledge of Docker, ECR, ECS task definitions/services, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks.
  • AWS Networking & Security: Proficiency in VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security best practices.
  • Incident Response & Observability: Skilled in troubleshooting using logs, metrics, alarms, deployment history, root cause analysis, rollback decisions, and operational runbooks.
  • Linux & Automation: Strong Linux and Git fundamentals with Bash/Python scripting for AWS CLI automation, CI/CD, and operational tooling.
  • AI Tooling Deployment: Hands-on experience integrating and operating AI services and APIs in production, including monitoring, reliability, and security practices for AI-powered features.
  • Developer Enablement: Ability to support multiple engineering teams, troubleshoot across infrastructure and application layers, document solutions, and enable secure self-service practices.


U.S. National Base Pay Range: $95,300 - $158,800. Geographic differentials may apply in some locations to better reflect local market rates.This job is eligible for an annual incentive bonus.
We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits. Click here to access benefits specific to your location.

About Relx Group

RELX Group is a global provider of information-based analytics and decision tools for professional and business customers. The company operates in four market segments: scientific, technical and medical; risk and business analytics; legal; and exhibitions. RELX's products and services include electronic databases, online information services, workflow tools, and print and digital books. The company was founded in 1993 and is headquartered in London, England.
Learn more about Relx Group
Size
33,500 employees
Market Cap
$53.1 billion
Industry
Net Income
$1.2 billion
Founded
2018
5 Year Trend
+1%
Revenue
$7.1 billion
NASDAQ

Similar Jobs

More Jobs at Relx Group

More Information Technology Jobs

Find similar Senior Site Reliability Engineer I jobs: