Job DescriptionThe Impact You Will Have in This RoleAs a
Principal Application Support Engineer, you will lead the reliability, resilience, and operational excellence of DTCC's mission-critical platforms that power the global financial markets. Operating with a strong
Site Reliability Engineering (SRE) mindset, you will drive production stability through automation, observability, proactive risk management, and continuous improvement.
Partnering across Engineering, Infrastructure, Cybersecurity, and Operations teams, you will champion platform resiliency, reduce operational complexity, accelerate incident resolution, and strengthen system performance at scale. As a senior technical leader, you will shape operational strategy, drive engineering best practices, and enhance the reliability of critical services that support DTCC's business and clients worldwide.
Your Primary Responsibilities:
- Provide technical leadership for the reliability, availability, and operational support of mission-critical applications and platforms.
- Lead major incident, problem, change, and production support processes, driving rapid resolution and long-term remediation.
- Partner with Engineering, Infrastructure, Cloud, and Operations teams to improve resiliency, observability, scalability, and performance.
- Drive operational excellence initiatives focused on automation, self-healing capabilities, alert optimization, and reduction of operational toil.
- Establish and monitor service reliability objectives, operational KPIs, and platform health metrics.
- Lead production readiness reviews, release management, disaster recovery testing, and business continuity planning.
- Strengthen operational risk management, regulatory compliance, audit readiness, and control effectiveness.
- Mentor engineers and foster a culture of accountability, collaboration, innovation, and continuous improvement.
- Serve as a trusted technical advisor during critical production events and strategic operational initiatives.
Qualifications - 12+ years of experience in Application Support, Production Support, Site Reliability Engineering (SRE), or Technology Operations.
- 3+ years of experience leading technical teams or large-scale operational initiatives.
- Bachelor's degree preferred or equivalent experience.
- Experience supporting highly available, business-critical applications within Financial Services or Capital Markets environments.
Talent Needed for Success - Deep expertise in Application Support Engineering, Site Reliability Engineering (SRE), and enterprise production operations.
- Strong experience leading major incident management, root cause analysis, and executive-level communications.
- Hands-on knowledge of Linux, Windows, AWS, OpenShift/Kubernetes, IBM MQ, and distributed application architectures.
- Experience with observability platforms such as Splunk, Grafana, and related monitoring technologies.
- Strong understanding of SQL/PLSQL, query optimization, middleware technologies, networking, Autosys, and enterprise application ecosystems.
- Experience driving automation through scripting, AI-enabled operational tooling, and workflow orchestration.
- Familiarity with ServiceNow and ITIL-based Incident, Problem, and Change Management practices.
- Exposure to mainframe-integrated environments, batch processing, and complex enterprise workflows.
- Exceptional communication, stakeholder management, and technical leadership skills.
- Capital Markets or Financial Services industry experience required.
The salary range is indicative for roles at the same level within DTCC across all US locations. Actual salary is determined based on the role, location, individual experience, skills, and other considerations.
Pay and Benefits: - Competitive compensation, including base pay and annual incentive
- Comprehensive health and life insurance and well-being benefits, based on location
- Pension / Retirement benefits
- Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
- DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee).