DTCC

Principal Resiliency Engineer

DTCC$130K — $160K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in distributed application design and implementation
  • Bachelor's degree in computer engineering or equivalent experience
  • 5+ years expertise in enterprise Java technologies and open standards
  • 5+ years in infrastructure, networking, middleware, and database architecture
  • 5+ years in highly available architecture and disaster recovery
  • 3+ years in containers and cloud-based solution delivery
  • Strong troubleshooting and performance analysis skills
  • Proficiency in Java, Python, Bash, SQL
  • Experience with CI/CD pipelines and automation frameworks

Responsibilities

  • Build and deploy prototype applications to demonstrate observability capabilities in hybrid environments
  • Conduct testing with simulated disruptions to validate observability standards
  • Collaborate with platform teams and external vendors for integrating observability
  • Produce documentation like runbooks and architectural patterns to support enterprise scalability
  • Provide feedback to IT Architecture to influence enterprise strategy
  • Evaluate proposed standards for observability visualization and metrics operations
  • Host sprint retrospectives and demos for stakeholder engagement

Benefits

  • Comprehensive health and wellness programs
  • Flexible work schedules and remote work options
  • Retirement savings plan with company contribution
  • Professional development and continuing education opportunities
  • Generous paid time off and holiday policies
Full Job Description
Job Description

The Impact you will have in this role:

We are seeking a Principal Reliability Engineer to join our Reliability Architecture team and play a pivotal role in shaping the future of enterprise observability. This hands-on engineering role is designed for a candidate who thrives in fast-paced, collaborative environments and is eager to influence architectural standards across the organization.

You will engineer prototype workloads that simulate real-world business applications, demonstrating how proposed resiliency and observability strategies and standards can be embedded into modern, cloud-native environments. Your work will directly support enterprise-wide adoption of our observability mandate and contribute to platform modernization and resiliency goals.

Primary Responsibilities:

  • Build and deploy prototype applications that showcase observability capabilities across hybrid environments (on-premises, AWS, Azure, SaaS).


  • Conduct rigorous testing with simulated disruptions, environmental failures, and performance scenarios to validate proposed observability standards.


  • Collaborate with platform teams, application owners, and external vendors to integrate observability into real workloads.


  • Produce runbooks, configuration guides, architectural patterns, reusable dashboards, and findings to support enterprise enablement and scale.


  • Provide feedback and technical insight to IT Architecture leadership and teams to influence enterprise strategy.


  • Conduct technical evaluations of proposed standards for observability visualization, distributed tracing, enterprise logging, metrics operations, open standards adoption, event correlation, and alerting notification.


  • Host sprint retrospectives and demos for engineering, business, and executive stakeholders to validate and promote adoption.


Qualifications:

  • Minimum of 8 years in distributed application design and implementation


  • Bachelor's degree in computer engineering or equivalent experience


Talents needed for success:

  • 5+ years in enterprise Java technologies and open standards


  • 5+ years in infrastructure, networking, middleware, and database architecture


  • 5+ years in highly available architecture and disaster recovery


  • 3+ years in containers and cloud-based solution delivery


  • Strong troubleshooting and performance analysis skills


  • Java, Python, Bash, SQL


  • CI/CD pipelines and automation frameworks (e.g., Jenkins, Selenium)


  • OpenTelemetry, distributed tracing, metrics generation, logging


  • Chaos engineering tools (e.g., Gremlin, AWS FIS)


  • AWS and Azure cloud environments


  • Infrastructure as Code (IaC), container orchestration, hybrid deployments


The salary range is indicative for roles at the same level within DTCC across all US locations. Actual salary is determined based on the role, location, individual experience, skills, and other considerations.

About DTCC

The Depository Trust & Clearing Corporation (DTCC) is a financial services company that provides clearing, settlement, and information services for the global financial industry. DTCC was founded in 1999 and is headquartered in New York City. The company operates through subsidiaries that provide services such as trade matching, risk management, and asset servicing. DTCC is owned by its users, which include broker-dealers, banks, and other financial institutions. The company is committed to reducing risk and increasing efficiency in the financial markets.
Learn more about DTCC
Size
4,000 employees
Industry
Founded
1973

Similar Jobs

More Jobs at DTCC

More Information Technology Jobs

Find similar Principal Resiliency Engineer jobs: