Site Reliability Engineering (SRE) Director is responsible for leading the strategy, operational discipline, and engineering practices that ensure the reliability, availability, scalability, and performance of technical and digital products and platforms. This role drives automation, observability, incident response maturity, and continuous improvement while building a SRE function that supports both enterprise continuity and innovation.
Responsibilities- Oversees the reliability and operational readiness of production systems, ensuring platforms can scale effectively and perform consistently under changing demand.
- Implements and continuously improves automated testing (e.g., QA, regression, performance, and reliability testing) within CI/CD pipelines to proactively identify defects, reduce production risk, and ensure consistent, high-quality releases.
- Develops and implements reliability goals, service level objectives, and performance expectations for critical digital products and platforms.
- Establishes and strengthens monitoring, logging, alerting, and observability practices to provide clear visibility into service health and performance.
- Owns the operational framework for incident management, change management, escalation, service restoration, root cause analysis, and post-incident review.
- Partners with digital teams to embed reliability, automation, and operational excellence across the software lifecycle and platforms.
Qualifications- Bachelor's degree or equivalent experience in Computer Science, Information Technology, Software Engineering, or a related field; advanced degree preferred.
- Minimum 8 years of progressive experience in Site Reliability Engineering, DevOps, cloud operations, infrastructure engineering, or a related field, including leadership of enterprise-scale reliability and operations functions.
- Demonstrated experience developing and executing SRE strategies, operational frameworks, and reliability programs that improve system availability, scalability, performance, and organizational resiliency.
- 5+ years in cloud-native architectures, distributed systems, infrastructure automation, CI/CD pipelines, and modern software delivery practices supporting mission-critical platforms, with experience establishing and governing Secure Software Development Lifecycle (SSDLC) practices that embed security, compliance, reliability, and risk management throughout the software delivery lifecycle.
AARP will not sponsor an employment visa for this position at this time.
Additional Requirements- Regular and reliable job attendance
- Effective verbal and written communication skills
- Exhibit respect and understanding of others to maintain professional relationships
- Independent judgement in evaluation options to make sound decisions
- In office/open office environment with the ability to work effectively surrounded by moderate noise
Hybrid Work EnvironmentAARP observes Mondays and Fridays as remote workdays, except for essential functions. Remote work can only be done within the United States and its territories.
Compensation and BenefitsAARP offers a competitive compensation and benefits package including a 401(k); 100% company-funded pension plan; health, dental, and vision plans; life insurance; paid time off to include company and individual holidays, vacation, sick, caregiving, and parental leave; performance-based and peer-based recognition and tuition reimbursement.