Site Reliability Engineer

Intercontinental Exchange Holdings, Inc.

$110K — $130K *
Finance & Insurance
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience
  • 3+ years in site reliability, production engineering, or software operations role
  • Strong technical skills with a history of delivering important work
  • Active involvement in training and mentoring team members
  • Ability to prioritize tasks and work independently without direct management guidance

Responsibilities

  • Employ advanced troubleshooting and root-cause analysis for platform services
  • Collaborate with Product and Engineering teams to deploy product releases
  • Work with Engineering leadership to build and evolve shared services
  • Design and implement proactive monitoring and self-healing automation
  • Resolve product defects and infrastructure issues with increasing independence
  • Implement automated tests and operational tooling across the SRE toolchain
  • Lead smaller projects and provide updates to management and stakeholders
  • Mentor junior engineers and contribute to team knowledge-sharing

Benefits

  • Opportunities for professional development and continued education
  • Autonomous work environment for significant impact on platform reliability
  • Involvement in AI-assisted automation and cutting-edge technology
  • Chance to collaborate across diverse teams and influence project direction
  • Mentoring roles to help foster growth within the engineering team
Full Job Description
Overview

Job Purpose

We are seeking a Site Reliability Engineer II to bring 3+ years of hands-on experience to our SRE team, operating with significant autonomy to improve platform reliability, drive automation, and mentor junior engineers. The ideal candidate contributes meaningfully to platform release cycles, leads smaller projects, and actively shapes the team's approach to observability, incident response, and service design in ICE's 24x7 production environment.

 

 

Responsibilities

  • Employ advanced troubleshooting and root-cause analysis to improve availability, performance, and security of IMT and platform services
  • Collaborate with Product and Engineering teams to plan and deploy product releases with operational rigor and quality gates
  • Work with Engineering leadership to build and evolve shared services meeting the requirements of platform and application teams
  • Design and implement proactive monitoring, alerting, trend analysis, and self-healing automation
  • Resolve product and service defects, infrastructure issues, and operational changes with increasing independence
  • Implement automated tests, automated deployments, and operational tooling across the SRE toolchain
  • Ensure services are designed with 24x7 availability and operational readiness and rigor
  • Lead smaller projects and provide status updates to management and stakeholders
  • Mentor SRE I engineers and contribute actively to team training and knowledge-sharing
  • Partner with application and platform teams to identify critical workflows and build automated health checks that run post-deployment and during incidents to accelerate root-cause identification
  • Design and build AI-assisted automated diagnosis jobs that correlate signals across monitoring and alerting platforms to reduce Mean Time to Resolution (MTTR) for production incidents
  • Build and maintain automation pipelines (e.g., Rundeck, Jenkins) that integrate with AI/LLM tooling to drive efficiency gains in observability, runbook execution, and incident triage
  • Develop and tune AWS CloudWatch metrics, alarms, and dashboards, instrument services using OpenTelemetry/Alloy, and build observability visualizations in Grafana; integrate alerting and event correlation workflows across PagerDuty, BigPanda, and Splunk to ensure timely, actionable incident notification

 

Knowledge and Experience

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience
  • 3+ years of experience in a site reliability, production engineering, or software operations role
  • Proven technical skills with strong personal initiative and consistent delivery of important work
  • Excellent teamwork with active involvement in training and mentoring
  • Ability to prioritize and execute without direct management guidance
  • Strong understanding of ICE Core Competencies

 

Preferred Knowledge and Experience

  • Experience in financial services technology, mortgage platforms, or exchange infrastructure
  • Familiarity with SRE principles including SLI, SLO, and error budget management
  • Exposure to Terraform, Chef, Ansible, or equivalent infrastructure automation frameworks
  • Hands-on experience with AWS observability services, CloudWatch, Grafana, OpenTelemetry/Alloy, Splunk, BigPanda, PagerDuty, and job orchestration/automation platforms such as Rundeck and Jenkins
  • Practical experience integrating AI/LLM-based tooling into operational workflows to automate diagnosis, reduce manual triage, and improve incident response efficiency

#LI-JM1

Similar Jobs

More Jobs at Intercontinental Exchange Holdings, Inc.

More Finance & Insurance Jobs

Find similar Site Reliability Engineer jobs: