Site Reliability Engineer, Lead
The Opportunity:As a Lead Site Reliability Engineer (SRE) on our team, you'll be responsible for ensuring the reliability, performance, scalability, and security of critical production systems and platforms. This role leads the design and implementation of observability, automation, incident response, and operational best practices across cloud and air-gapped environments, while partnering closely with DevOps, infrastructure, and security teams to improve system resilience and reduce operational risk. The Lead SRE also drives root cause analysis, capacity planning, reliability standards, and continuous improvement initiatives to support highly available, efficient, and scalable services. This is your chance to further your skills in cloud infrastructure and technologies while continuing to grow your SRE experience. You'll build and support a reliable site for the environment in order to meet the development and maintenance requirements of systems and platforms. Work with the development and operation teams to evaluate the health, stability and reliability of systems and platforms. Design and develop technical tools to debug problems that occur in the deployment of applications, within specific platforms and systems.
Join our efforts to strengthen our security posture and safeguard national interests.
Join us. The world can't wait.
You Have: - 8+ years of experience with monitoring, logging, and observability platforms, such as Prometheus, Grafana, and ELK stack
- 8+ years of experience with Linux systems administration and networking fundamentals within AWS
- Experience with Python scripting and automation
- Experience with Infrastructure as Code using Terraform and Terragrunt
- Knowledge of Kubernetes administration, troubleshooting, and operations.
- TS/SCI clearance with a polygraph
- Bachelor's degree and 8+ years of experience in Site Reliability Engineering, DevOps Engineering, or Platform Engineering, or 12+ years of experience in Site Reliability Engineering, DevOps Engineering, or Platform Engineering in lieu of a degree
- Ability to obtain a Security+ CE, SSCP, CCNA-Security, or GSEC Certification within 6 months of start date
Nice If You Have: - Experience with deploying and managing OpenTelemetry.
- Experience with AWS CloudWatch, AWS EKS, and related AWS services
- Experience managing Kubernetes environments through Rancher
- Experience implementing SRE practices such as SLOs, SLIs, error budgets, and incident management
- Experience with Jenkins, Git, Docker, Kubernetes, Nessus, JIRA, and Confluence
- Knowledge of distributed systems, microservices architectures, and containerized workloads
- Knowledge of NIST 800-53 and NIST-190
- Master's degree in a relevant field
- Security+ CE, SSCP, CCNA-Security, or GSEC Certification
Clearance: Applicants selected will be subject to a security investigation and may need to meet eligibility requirements for access to classified information; TS/SCI clearance with polygraph is required.
CompensationSalary at Booz Allen is determined by various factors, including but not limited to location, the individual's particular combination of education, knowledge, skills, competencies, and experience, as well as contract-specific affordability and organizational requirements. The projected compensation range for this position is $99,000.00 to $225,000.00 (annualized USD). The estimate displayed represents the typical salary range for this position and is just one component of Booz Allen's total compensation package for employees. This posting will close within 90 days from the Posting Date.
Work ModelOur people-first culture prioritizes the benefits of collaboration, whether it occurs in person or virtually. To support engagement and effective communication, employees working virtually are generally expected to have their cameras on during meetings.
- Remote: If this position is listed as remote, there may still be occasions when you are required to work in person at a Booz Allen or customer facility.
- Hybrid: If this position is listed as hybrid, you will be expected to work from a Booz Allen facility frequently, in alignment with leadership expectations and the needs of the role. You may also be required to work from or visit a customer facility.
- Onsite: If this position is listed as onsite, work will primarily be performed at a Booz Allen office or customer facility, where employees will collaborate directly with colleagues and customers as required by the role.