Job DescriptionThe impact you will have in this role:
As a Senior Application Support Engineer (SRE), you will play a critical role in ensuring the stability, reliability, and performance of mission-critical applications at DTCC.
This role goes beyond traditional support-focusing on Site Reliability Engineering principles, proactive system improvement, and operational excellence. You will partner closely with development, infrastructure, and global operations teams to enhance system resilience, reduce operational toil, and drive continuous improvement across the platform.
Your Primary Responsibilities:
- Act as a Lead Application Support Engineer with SRE responsibilities, partnering with engineering and infrastructure teams to improve system reliability, resilience, and observability
- Lead the resolution of critical production incidents, providing clear impact analysis, root cause identification, and preventive actions
- Own and drive incident, problem, and major incident management, including post-incident reviews and continuous improvement
- Proactively identify reliability risks and implement solutions to prevent recurrence and reduce operational toil
- Develop, maintain, and enhance runbooks, knowledge articles, and operational documentation
- Execute and support release, change, and deployment activities, including production releases and vendor upgrades
- Support and participate in Disaster Recovery (DR) testing, execution, and audit readiness
- Drive automation and alert optimization initiatives to improve efficiency and reduce noise
- Embed risk, control, and reliability best practices into day-to-day operations
- Collaborate with global teams to ensure high availability and operational excellence across systems
**NOTE: The Primary Responsibilities of this role are not limited to the details above. **
Qualifications: - 6+ years of experience in application support, SRE, or production engineering
- Bachelor's degree preferred or equivalent experience
Required Skills• Strong hands-on experience in Application Production Support with SRE mindset focused on reliability, observability, production stability, resiliency, and incident prevention.
• Extensive experience supporting Linux and Windows environments, encompassing process inspection, log analysis, advanced troubleshooting, performance diagnostics, and system optimization.
• Strong scripting and automation expertise across Bash, Shell Scripting, Python, Ruby, Perl, and JavaScript.
• Hands-on experience with enterprise monitoring, observability, and analytics platforms such as Dynatrace, Splunk, Grafana, and Selenium.
• Deep technical expertise in SQL databases such as Oracle, Snowflake, PostgreSQL, or similar technologies, with proven ability to perform data analysis, troubleshoot production issues, conduct root cause investigations, and optimize query performance.
• Strong experience with ITSM and operational management platforms such as ServiceNow and Jira, supporting Incident, Problem, Change, and Major Incident Management processes.
• Hands-on experience supporting enterprise Messaging and Queueing Technologies such as IBM MQ, Oracle AQ, ActiveMQ, RabbitMQ, and Kafka.
• Strong experience working with enterprise job scheduling platforms, particularly Autosys.
• Working knowledge of Cloud and Container Technologies such as OpenShift, AWS (EC2, S3, Lambda, SQS, IAM Roles), Amazon RDS Aurora, PostgreSQL, and cloud-native operational support practices.
• Familiarity with Mainframe technologies and troubleshooting concepts across COBOL, JCL, DB2, DB2 Stored Procedures, CICS, SPUFI, and File-AID.
• Strong understanding of security, risk, and operational controls, including certificate management, password management, access controls, and security best practices.
• Exposure to Capital Markets and Financial Services environments with experience supporting mission-critical, high-availability, and highly regulated systems.
• Knowledge of Artificial Intelligence concepts and their practical application within Production Support, Operational Engineering, and Service Reliability disciplines.
• Demonstrated ability to communicate effectively across all organizational levels with exceptional verbal, written, and stakeholder management skills.
• Proven leadership qualities with a strong sense of ownership, accountability, urgency, and operational excellence in fast-paced production environments.
• Ability to collaborate effectively with global, geographically distributed teams, driving outcomes across technology, infrastructure, operations, and business functions.
• Proactive and continuous improvement mindset with a passion for automation, innovation, operational efficiency, reliability engineering, and reducing operational toil through engineering-led solutions.
• Strong analytical, problem-solving, and critical-thinking capabilities with the ability to navigate complex technical challenges, drive root-cause resolution, and deliver sustainable long-term solutions.
• Demonstrated ability to thrive in high-pressure environments while maintaining focus, composure, and an unwavering commitment to service availability, customer experience, and operational resilience.
The salary range is indicative for roles at the same level within DTCC across all US locations. Actual salary is determined based on the role, location, individual experience, skills, and other considerations.
About the TeamServes as a dedicated technology resource for advancing DTCC's business opportunities and providing industry thought leadership for leveraging new technology. The goal of this new department is to partner internally with IT, our business and regulatory divisions and externally with clients, regulators, and fintech vendors, to help build new platforms and business models to advance DTCC's mission to support the financial markets.