DTCC

Principal Site Reliability Engineer (Cloud, Observability & Automation)

DTCC$150K — $180K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience.
  • 8+ years in Site Reliability Engineering, Production Engineering, DevOps, or related fields.
  • Strong hands-on experience with AWS and cloud-native architectures.
  • Proficiency in programming languages such as Python, Java, or Go.
  • Strong Linux/Unix systems administration and troubleshooting capabilities.
  • Expertise in observability and monitoring platforms including Splunk and Grafana.
  • Excellent communication and stakeholder management skills.

Responsibilities

  • Drive reliability and scalability for critical enterprise applications.
  • Design and implement observability solutions using monitoring platforms.
  • Define and manage service level indicators (SLIs) and operational KPIs.
  • Lead major incident response and root cause analysis initiatives.
  • Build self-healing capabilities and automation solutions using modern technologies.
  • Partner with cross-functional teams to embed SRE best practices in the software lifecycle.
  • Identify operational risks and implement reliability improvements.

Benefits

  • Opportunities for career development and advancement.
  • Access to innovative technologies and methodologies.
  • Collaborative work environment with cross-functional teams.
  • Engagement with a mission-critical technology ecosystem.
Full Job Description
Job Description

The Impact You Will Have in This Role

The Enterprise Application Support (EAS) team supports critical applications across the ITP and ECS business lines, ensuring the reliability, scalability, and performance of enterprise platforms.

As a Principal Site Reliability Engineer (SRE), you will drive operational excellence across mission-critical systems. You will lead reliability initiatives, champion observability and automation, drive major incident response, and partner with engineering, infrastructure, and security teams to build resilient, highly available applications.

This is a hands-on technical leadership role focused on improving system performance, reducing operational risk, accelerating recovery, and advancing SRE best practices through modern cloud, observability, automation, and AI-powered technologies.

Your Primary Responsibilities:
  • Drive reliability, scalability, resiliency, and operational excellence across critical enterprise applications.
  • Design and implement observability solutions using Splunk, Grafana, Dynatrace, ITSI, and related monitoring platforms.
  • Define and manage SLIs, SLOs, dashboards, alerts, and operational KPIs.
  • Lead major incident response, root cause analysis, and continuous service improvement initiatives.
  • Build automation, self-healing capabilities, and AI-assisted operational solutions using Python, Java, Amazon Q, Kiro, and related technologies.
  • Partner with development, infrastructure, cloud, security, and application teams to embed SRE best practices throughout the software development lifecycle.
  • Drive operational readiness, capacity planning, performance optimization, disaster recovery, and resiliency initiatives.
  • Identify operational risks and deliver strategic reliability improvements across the technology ecosystem.
  • Collaborate with technical and business stakeholders to improve service reliability and operational outcomes.

Qualifications
  • Bachelor's degree in Computer Science, Engineering, or equivalent experience.
  • 8+ years of experience in Site Reliability Engineering, Production Engineering, DevOps, Application Support Engineering, or related disciplines.

Talent Needed for Success
  • Strong hands-on experience with AWS and cloud-native architectures.
  • Proficiency in Python, Java, Go, or similar programming languages.
  • Strong Linux/Unix systems administration and troubleshooting experience.
  • Expertise in observability and monitoring platforms including Splunk, Grafana, Dynatrace, and ITSI.
  • Experience leading major incident management and root cause investigations in complex production environments.
  • Strong understanding of distributed systems, resiliency engineering, performance tuning, automation, and operational excellence.
  • Excellent communication and stakeholder management skills with the ability to influence technical and business partners.

Preferred Qualifications
  • Experience with AI-assisted engineering tools such as Amazon Q, Kiro, or similar technologies.
  • Experience designing and measuring SLOs, SLIs, and operational KPIs.
  • Experience supporting large-scale enterprise applications in financial services or other highly regulated environments.

The salary range is indicative for roles at the same level within DTCC across all US locations. Actual salary is determined based on the role, location, individual experience, skills, and other considerations.

About DTCC

The Depository Trust & Clearing Corporation (DTCC) is a financial services company that provides clearing, settlement, and information services for the global financial industry. DTCC was founded in 1999 and is headquartered in New York City. The company operates through subsidiaries that provide services such as trade matching, risk management, and asset servicing. DTCC is owned by its users, which include broker-dealers, banks, and other financial institutions. The company is committed to reducing risk and increasing efficiency in the financial markets.
Learn more about DTCC
Size
4,000 employees
Industry
Founded
1973

Similar Jobs

More Jobs at DTCC

More Information Technology Jobs

Find similar Principal Site Reliability Engineer (Cloud, Observability & Automation) jobs: