Relx Group

Senior Site Reliability Engineer II

Relx Group • $104K — $174K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years experience in Site Reliability Engineering, Systems Engineering, DevOps, or related fields.
  • Bachelor's degree in Engineering, Computer Science, IT, or equivalent experience.
  • Experience supporting highly available production systems.
  • Proven leadership in incident reviews and root-cause analysis.
  • Collaboration experience across infrastructure, application, security, and operations teams.
  • Strong analytical, organizational, and communication skills.
  • Ability to manage multiple priorities in a fast-paced environment.

Responsibilities

  • Lead incident response and conduct postmortems to identify system issues.
  • Analyze and mitigate reliability, availability, performance, and security risks in production systems.
  • Develop and track corrective actions for system vulnerabilities.
  • Coordinate with stakeholders for timely incident resolution and gap closure.
  • Respond to system alerts and manage operational exceptions seamlessly.
  • Input into project plans and operational readiness for multiple systems.
  • Support the execution and documentation of operational tasks and changes.

Benefits

  • Comprehensive health benefits package for medical, dental, and vision.
  • 401(k) retirement plan with employer matching and employee stock purchase options.
  • Wellness platform with incentives and subscriptions for mental health applications.
  • Short- and long-term disability and life insurance coverage.
  • Family benefits including bonding leave and adoption assistance.
  • Various spending accounts for health, dependent care, and commuting.
  • Additional paid leave for participation in employee resource groups or volunteering.
Full Job Description
About the Role:

The SRE role is responsible for improving the reliability, availability, performance, and operational quality of production systems. This role provides technical input into project plans, schedules, methodologies, and operational strategies across multiple system environments.

Job Functions
  • Lead and participate in incident response, postmortems, root-cause analysis, and gap assessments.
  • Identify reliability, availability, performance, security, and operational risks across production environments.
  • Develop, prioritize, and track corrective and preventive actions through completion.
  • Follow up with engineering, development, security, support, and business stakeholders to ensure timely resolution of incidents and identified gaps.
  • Respond to system-management alerts and operational exceptions within assigned enterprise systems and product offerings.
  • Provide technical input into project plans, schedules, implementation methodologies, and operational readiness activities.
  • Support the triage, planning, execution, documentation, and closure of changes, service requests, and operational tasks.
  • Lead or contribute to Operations Team projects involving cloud, on-premises infrastructure, security, Kubernetes, automation, monitoring, and system modernization.
  • Improve production quality and availability by creating new operational capabilities and remediating weaknesses in existing systems and processes.


Qualifications
  • 5+ years of experience in Site Reliability Engineering, Systems Engineering, DevOps, Infrastructure Engineering, or a related field.
  • Bachelor's degree in Engineering, Computer Science, Information Technology, or equivalent professional experience.
  • Demonstrated experience supporting highly available production systems.
  • Experience leading incident reviews, postmortems, root-cause analysis, and remediation planning.
  • Experience working across infrastructure, application, security, and operations teams.
  • Strong problem-solving, analytical, organizational, and communication skills.
  • Ability to manage multiple priorities and drive work to completion in a fast-paced operational environment.


Technical Skills

  • Strong experience in Site Reliability Engineering (SRE),production operations, and IT service management processes including incident, problem, change, and service request management.
  • Hands-on expertise with cloud and on-premises infrastructure, Kubernetes, containerized workloads, virtualization, and distributed systems.
  • Advanced knowledge of Linux/UNIX and Windows environments, storage and file systems, including installation, configuration, troubleshooting, lifecycle management, backup, disaster recovery, and business continuity.
  • Experience with monitoring, alerting, logging, observability, and performance analysis, including the ability to analyze system diagnostics, logs, traces, resource utilization, and operational metrics.
  • Strong automation and infrastructure engineering skills, including Infrastructure as Code (IaC), configuration management, scripting (Python, Shell, PowerShell), system provisioning, deployments, remediation, and security risk mitigation.


Accountabilities
  • Monitor assigned environments, respond to alerts and incidents, diagnose system and performance issues, and coordinate escalation and recovery.
  • Track remediation activities and stakeholder commitments through completion to improve production quality, reliability, and availability.
  • Design and maintain automation, scripts, integrations, runbooks, and workflows for provisioning, health checks, deployments, remediation, and routine operations.Install, configure, troubleshoot, and support hardware, software, storage, network, cloud, Kubernetes, and other infrastructure services.
  • Establish logging, monitoring, alerting, metrics, and tracing standards; improve alert quality by reducing noise and ensuring alerts are actionable.Build dashboards and visualizations that communicate system health, availability, performance, capacity, service-level objectives, and incident trends.
  • Develop and maintain recovery procedures and participate in disaster-recovery, resilience, and business-continuity exercises.Partner with development, operations, security, support teams, vendors, and stakeholders to coordinate work, resolve issues, and meet delivery commitments.
  • Lead or contribute to Operations Team projects from planning and implementation through documentation, transition to support, and closure.
  • Plan, risk-assess, obtain approval for, implement, document, and close changes, service requests, and operational tasks.
  • Review and improve technical procedures, scripts, automation, and operational documentation while providing guidance to less-experienced team members.


Working for you:

We know that your wellbeing and happiness are key to a long and successful career. These are some of the benefits we are delighted to offer:

  • Health Benefits: Comprehensive, multi-carrier program for medical, dental and vision benefits
  • Retirement Benefits: 401(k) with match and an Employee Share Purchase Plan
  • Wellbeing: Wellness platform with incentives, Headspace app subscription, Employee Assistance and Time-off Programs
  • Short-and-Long Term Disability, Life and Accidental Death Insurance, Critical Illness, and Hospital Indemnity
  • Family Benefits, including bonding and family care leaves, adoption and surrogacy benefits
  • Health Savings, Health Care, Dependent Care and Commuter Spending Accounts
  • In addition to annual Paid Time Off, we offer up to two days of paid leave each to participate in Employee Resource Groups and to volunteer with your charity of choice


U.S. National Base Pay Range: $104,900 - $174,700. Geographic differentials may apply in some locations to better reflect local market rates.This job is eligible for an annual incentive bonus.
We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits. Click here to access benefits specific to your location.

About Relx Group

RELX Group is a global provider of information-based analytics and decision tools for professional and business customers. The company operates in four market segments: scientific, technical and medical; risk and business analytics; legal; and exhibitions. RELX's products and services include electronic databases, online information services, workflow tools, and print and digital books. The company was founded in 1993 and is headquartered in London, England.
Learn more about Relx Group
Size
33,500 employees
Market Cap
$53.1 billion
Industry
Net Income
$1.2 billion
Founded
2018
5 Year Trend
+1%
Revenue
$7.1 billion
NASDAQ

Similar Jobs

More Jobs at Relx Group

More Information Technology Jobs

Find similar Senior Site Reliability Engineer II jobs: