As the Resiliency Engineer, you help design and evolve QTS’s Infrastructure Resiliency capability – working across the architecture for backup, disaster recovery orchestration, and cyber-resilient infrastructure. You help set the design standards, integration patterns, and technology roadmap that the broader engineering team implements and operates against.
What You Will Do:
- Own the end-to-end architecture for QTS’s resiliency platform portfolio across backup (Cohesity, FortKnox, Azure, AWS), SaaS protection (GitProtect), identity DR (Semperis), and DR orchestration.
- Define and maintain design standards, reference architectures, and integration patterns for all recovery and resilience platforms.
- Own the Infrastructure Resiliency roadmap; refresh it on at least an annual cadence and as the threat, vendor, and infrastructure landscape shifts.
- Lead platform technology selection and vendor evaluation – proofs of concept, capability assessments, and architectural recommendations.
- Design the target-state recovery architecture and evolve it continuously as QTS’s infrastructure footprint, cloud estate, and site portfolio grow.
- Own capacity and growth planning across the resiliency platform portfolio – forecasting storage, compute, and licensing needs and feeding those projections into refresh, expansion, and budget planning.
- Establish integration and architecture-review standards for every new system entering resiliency scope, ensuring recovery is designed in – not bolted on.
- Partner with the broader engineering team, InfraOps, enterprise architecture, and the CISO organization on recovery topology, dependency mapping, and cyber-resilient design.
- Other duties as assigned.
What You Will Need to Be Successful:
- Bachelor’s degree in an IT-related field (Information Technology, Computer Science, Cybersecurity, Information Systems) or equivalent practical experience.
- US Citizenship required.
- 5+ years in enterprise infrastructure, data protection, or disaster recovery, including hands-on platform design experience.
- Demonstrated experience designing or contributing to recovery, backup, or DR platform architecture – technology selection, design standards, and integration across multiple platforms.
- Strong command of backup and recovery architecture, DR orchestration, and cloud (Azure/AWS) recovery design.
- Ability to translate recovery objectives (RTO/RPO) and business risk into sound, implementable platform architecture.
- Strong technical documentation, roadmap ownership, and the ability to influence engineering and infrastructure teams without direct authority.
Other Key Skills:
- Ability to define reference architectures and design standards across compute (VMware, Hyper-V, cloud), storage/SAN/NAS, replication, and data protection platforms.
- Familiarity with immutable, air-gapped, and clean-room recovery architecture patterns.
- Understanding of identity/DNS/network control planes (AD, Azure AD) and their impact on recovery topology.
- Vendor evaluation and technology-roadmap ownership in a fast-moving data center environment.
- Capacity and growth forecasting for storage, compute, and licensing.
- Regulatory and audit-driven architecture familiarity (CMMC, NIS2, DORA, SOC 2, ISO 27001, FedRAMP).
- Certifications such as TOGAF, CISSP, or platform architecture certifications (Cohesity, Azure/AWS) preferred.