TD Bank

Senior Technology Resilience and Availability Management Analyst

TD Bank • $96K — $136K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8-12 years in technology operations, site reliability engineering, or related fields; experience in financial services preferred.
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent experience.
  • Proven knowledge of high-availability design patterns and recovery metrics.
  • Hands-on experience with disaster recovery planning and exercises.
  • Familiarity with compliance frameworks such as ITIL, ISO 22301, and cybersecurity standards.

Responsibilities

  • Lead assessments to identify vulnerabilities and recovery gaps in technology infrastructure.
  • Design and improve systems for high-availability and disaster recovery initiatives.
  • Define and monitor resilience objectives like RTO and RPO.
  • Conduct impact analysis to align tech capabilities with business resilience needs.
  • Oversee tests and drills for failover and recovery scenarios.
  • Evaluate backup and cyber-recovery capabilities for robustness.
  • Embed resilience requirements into various technology operations and software lifecycles.

Benefits

  • Opportunities for skill development and career progression.
  • Access to discretionary variable compensation awards based on performance.
  • Commitment to fair and equitable compensation practices.
  • Supportive work environment with a focus on technology resilience and operational excellence.
Full Job Description

Work Location:

Toronto, Ontario, Canada

Hours:

37.5

Line of Business:

Technology Solutions

Pay Details:

$96,900 - $136,800 CAD

This role is eligible for a discretionary variable compensation award that considers business and individual performance.

TD is committed to providing fair and equitable compensation opportunities to all colleagues. Growth opportunities and skill development are defining features of the colleague experience at TD. Our compensation policies and practices have been designed to allow colleagues to progress through the salary range over time as they progress in their role. The base pay actually offered may vary based upon the candidate's skills and experience, job-related knowledge, geographic location, and other specific business and organizational needs.

As a candidate, you are encouraged to ask compensation related questions and have an open dialogue with your recruiter who can provide you more specific details for this role.

Job Description:


Role Summary


This first-line role supports the design, governance, assessment, and continuous improvement of Technology Resilience and Availability Management capabilities. The position works across application, infrastructure, cloud, cyber, risk, audit, and business teams to strengthen high availability, disaster recovery, backup and cyber recovery, capacity management, and operational resilience for critical technology services. The successful candidate combines strong technical depth with the ability to influence stakeholders, produce defensible evidence, and drive remediation in a complex regulated environment.



Job Responsibilities


  • Lead resilience and availability assessments across applications, infrastructure, data, network, cloud, and third-party services to identify vulnerabilities, single points of failure, recovery gaps, and control weaknesses.
  • Design, assess, and improve high-availability and recovery patterns, including active-active architectures, clustering, load balancing, multi-zone or multi-region deployment, automated failover, and resilient dependency design.
  • Define, validate, and monitor resilience objectives and measures, including Recovery Time Objective (RTO), Recovery Point Objective (RPO), Maximum Tolerable Downtime (MTD), service-level objectives and indicators, availability targets, and capacity.
  • Lead business and application impact analysis, dependency mapping, critical service mapping, and recovery prioritization to align technology capabilities with business resilience requirements.
  • Plan and oversee high-availability and failover exercises, recovery-from-backup tests, tabletop scenarios, extended-duration testing, and appropriate failure-injection or chaos-testing practices.
  • Assess backup, restore, and cyber-recovery capabilities, including immutable or isolated backups, point-in-time recovery, clean-room recovery, and ransomware recovery scenarios.
  • Embed resilience and availability requirements into technology architecture, service operations, capacity management, change management, configuration management, incident and problem management, and the software development lifecycle.
  • Use monitoring and observability data to identify availability, performance, capacity, and recovery risks, and translate findings into prioritized remediation actions and measurable improvements.
  • Coordinate remediation activities, track risks and actions to closure, and provide clear reporting on capability maturity, test outcomes, control effectiveness, and residual risk.
  • Produce organized, traceable, and defensible evidence for internal audit, regulatory examinations, senior management, and board or risk committee reporting.
  • Serve as a trusted resilience advisor and central point of coordination across engineering, application owners, infrastructure, cyber, business continuity, technology risk, third-party risk, and operational resilience teams.
  • Monitor emerging technology risks, regulatory expectations, cyber threats, and industry practices, and recommend practical enhancements to resilience standards, procedures, controls, and testing methods.

Job Requirements


 

Resilience, Availability and Recovery


  • Demonstrated knowledge of high-availability design patterns, failover strategies, fault tolerance, redundancy, and elimination of single points of failure.
  • Experience establishing and assessing RTO, RPO, MTD, availability objectives, SLOs, and recovery or availability metrics.
  • Experience with business or application impact analysis, critical service mapping, technology dependency mapping, and recovery sequencing.
  • Hands-on experience developing DR plans and runbooks and coordinating technical recovery exercises, failover tests, tabletop exercises, and end-to-end recovery validation.
  • Knowledge of enterprise backup and restore, immutable or isolated backup, point-in-time recovery, and cyber-recovery concepts.
  • Knowledge of capacity management, performance monitoring, utilization forecasting, and reporting

 

Platforms, Engineering and Tooling


  • Experience with cloud resilience capabilities in AWS, Microsoft Azure, and/or Google Cloud, including multi-region architecture, traffic management, native backup, and disaster recovery services.
  • Understanding of on-premises and hybrid technology, including VMware, SAN/NAS storage, Windows, Linux, Active Directory, DNS, and network dependencies.
  • Working knowledge of container and Kubernetes high-availability patterns in cloud-native environments.
  • Knowledge of database resilience methods for platforms such as Oracle, Microsoft SQL Server, and PostgreSQL, including replication, clustering, backup, restore, and point-in-time recovery.
  • Experience using observability and monitoring platforms such as Splunk, Dynatrace, Datadog, or equivalent tools to assess availability, capacity, performance, and recovery outcomes.
  • Experience with automation and infrastructure-as-code tools such as Python, or PowerShell.
  • Experience with ServiceNow capabilities, including CMDB, incident, problem, change, and related technology risk or control workflows, is strongly preferred.
  • Proficiency with Microsoft Word, Excel, and PowerPoint for analysis, evidence management, executive reporting, and program documentation.

 

Risk, Control and Regulatory


  • Strong understanding of technology risk, operational resilience, disaster recovery, business continuity, and control assessment in a regulated environment.
  • Working knowledge of relevant frameworks and guidance, including FFIEC Business Continuity Management expectations, OSFI/OCC and Federal Reserve operational-resilience guidance, NIST Cybersecurity Framework, and ITIL practices.
  • Experience mapping technology applications and dependencies to important or critical business services.
  • Knowledge of third-party resilience, including concentration risk, critical technology service provider testing, contingency planning, and exit strategies.
  • Experience preparing evidence and written responses for internal audit, regulators, risk committees, and senior executives.
  • Ability to apply major incident, problem, change, and post-incident review practices to improve resilience and reduce recurring disruption.

 


 

Competencies


  • Advanced stakeholder management and relationship-building skills across engineering, application, infrastructure, cyber, risk, audit, and business teams.
  • Clear and concise written and verbal communication, including executive-ready updates, root-cause analyses, postmortems, board or risk materials, and regulator-ready documentation.
  • Ability to influence without direct authority and drive adoption of standards, remediation commitments, and sustainable process improvements across a matrix organization.
  • Calm, structured leadership during incidents, recovery events, testing exercises, and periods of heightened scrutiny.
  • Strong analytical problem-solving, including root-cause analysis, dependency analysis, scenario analysis, risk assessment, and gap-to-control mapping.
  • Excellent organization and evidence-management discipline, with the ability to manage multiple priorities, deadlines, and stakeholders without compromising quality.
  • Sound judgment under ambiguity and the ability to design and assess severe but plausible scenarios rather than relying only on happy-path recovery assumptions.
  • Collaborative mindset suited to global, matrixed, and follow-the-sun operating models.
  • Continuous-improvement orientation focused on reducing outages, improving control effectiveness, increasing test fidelity, automating repeatable work, and closing findings sustainably.

 


Experience and Education


  • Typically 8 to 12 years of relevant experience in technology operations, site reliability engineering, infrastructure, disaster recovery, business continuity, availability or capacity management, cyber recovery, or technology risk; experience in financial services or another regulated industry is strongly preferred.
  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline, or equivalent practical experience.
  • Experience leading complex resilience initiatives, technical assessments, remediation programs, or cross-functional testing activities in a large enterprise environment.


Preferred Qualifications


  • ITIL 4, CBCI/MBCI, ISO 22301, CISSP, CISA, CRISC, cloud associate or professional certification, or relevant site reliability engineering training.
  • Experience applying SRE practices to regulated workloads, including SLOs, error budgets, automation, and toil reduction.
  • Experience with cyber-resilience and ransomware-recovery capabilities, including isolated recovery environments and clean-room restoration.
  • Exposure to mainframe, payments, core-banking, or other highly critical enterprise platforms.
  • Experience supporting regulatory examinations, internal audit reviews, or formal remediation programs related to technology resilience, availability, recovery, or business continuity.
  • Experience integrating resilience requirements and control gates into Agile, DevOps, architecture review, change, and software delivery processes.

 

About TD Bank

TD Securities offers a range of advisory and capital market services to its clients. The company's range of services includes research, investment banking, capital markets, and global transaction banking. Research consists of commodity and equity research. Investment banking consists of mergers, acquisitions, industry expertise, and credit origination. Global transaction banking consists of trade finance, cash management, and correspondent banking. TD Securities was founded in 1855 and is based in Ontario.

TD Bank Careers

Join the vibrant team at TD Bank, one of North America's leading financial services organizations, where innovation, leadership, and growth go hand in hand. At TD Bank, we are committed to fostering a culture of diversity and inclusion, making it an ideal place for ambitious professionals to thrive. Work You’ll Do At TD Bank, your professional journey is bolstered by a robust support system. From your first interview to every career milestone, you will find opportunities for growth and leadership. Our team is dedicated to helping you develop the skills necessary for success in the ever-evolving financial sector. TD Bank offers a variety of job opportunities across multiple fields, from customer service to investment banking. Each position at TD Bank is a chance to contribute to our culture of innovation and exceptional client service. Internship Programs Kickstart your career with a TD Bank internship. Our programs provide invaluable industry exposure and hands-on experience, making them a perfect starting point for students and recent graduates eager to make their mark in the banking industry. Interns at TD Bank enjoy the unique opportunity to work alongside seasoned professionals, gaining insights that are crucial for future employment. Benefits and Growth TD Bank is deeply committed to the well-being and continuous growth of our team members. We offer competitive benefits packages that cover health, finance, and family care. Our employees enjoy comprehensive health insurance, retirement plans, and generous paid time off, among other perks. Moreover, TD Bank encourages professional development through various training programs, including leadership development and diversity training. These initiatives ensure that our team remains at the forefront of industry standards and best practices. Join Our Team Explore the numerous career paths available at TD Bank and discover how your skills and interests align with our mission. We are actively hiring and continually looking for talented individuals who are passionate about banking and customer service. Networking and Professional Development At TD Bank, we believe in the power of networking and collaboration. Our employees have access to a wide range of networking events, workshops, and seminars that promote career development and professional growth. These platforms not only enhance your professional skills but also expand your industry connections. Stay Connected Keep up to date with the latest at TD Bank Careers by subscribing to our job alert emails. Tailor your subscription to match your career preferences and get the latest news, insider tips, and job opportunities delivered straight to your inbox. Explore job opportunities at TD Bank and be part of a team that values hard work, creativity, and a diverse workplace culture. Your next great career move is just a click away. SEARCH TD BANK JOBS Join us at TD Bank and let your ambition lead you to a rewarding career filled with opportunities to learn, grow, and innovate.
Learn more about TD Bank
Size
90,000 employees
Market Cap
$117.9 billion
Industry
Net Income
-$6.9 million
5 Year Trend
+6.6%

Similar Jobs

More Jobs at TD Bank

More Information Technology Jobs

Find similar Senior Technology Resilience and Availability Management Analyst jobs: