Lead Site Reliability Engineer

Peraton

• $112K — $179K *
US-AnywhereRemote in United States
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor’s degree with 8–10 years of relevant experience in SRE, DevOps, cloud, or infrastructure; or 12 years of experience with a high school diploma.
  • Expert knowledge of AWS services including compute, networking, storage, and serverless components.
  • Strong experience with Infrastructure as Code (Terraform, CloudFormation) and automation principles.
  • Experience in building CI/CD pipelines using GitHub actions or similar tools.
  • Deep understanding of Kubernetes and Docker-based deployments.
  • Proven experience in validating disaster recovery processes and executing system rebuilds.
  • Expertise in building monitoring tools and observability solutions using CloudWatch, Datadog or similar tools.
  • Proficiency in programming/scripting languages such as Python, Java, C#, or Go.

Responsibilities

  • Support full lifecycle platform portability and disaster recovery drill execution.
  • Execute infrastructure level drill activities to confirm platform rebuild capabilities within a 48-hour recovery target.
  • Verify data completeness and accuracy during drill exercises and document remediation recommendations.
  • Identify gaps in infrastructure and deployment processes, driving corrective actions with engineering teams.
  • Design and implement automated Infrastructure as Code workflows and support CI/CD pipelines.
  • Manage and optimize Kubernetes clusters and ensure workload reliability.
  • Build observability solutions to enhance service reliability and incident response.
  • Develop scripts and automation tools to improve operational processes.

Benefits

  • Public Trust clearance support.
  • Opportunities for certification and professional development.
  • Participate in managing high-visibility, mission-critical programs.
Full Job Description
Responsibilities

Peraton is seeking a Lead Site Reliability Engineer to join our team of qualified, diverse individuals. The ideal candidate will play a critical role in ensuring the reliability, resilience, and recoverability of mission essential platforms by leading infrastructure level disaster recovery drills, platform rebuild validation, and automated deployment processes. This engineer will partner closely with cross functional teams to maintain and enhance complex cloud-based environments, integrate modern automation solutions, and support largescale modernization and continuity efforts across high visibility programs.

 

Responsibilities

The Lead Site Reliability Engineer’s responsibilities shall include, but are not limited to:

  • Supporting full lifecycle platform portability and disaster recovery (DR) drill execution, including validation of platform rebuild procedures and DR playbooks.
  • Executing infrastructure level drill activities to ensure the platform can be fully rebuilt within the 48hour recovery target.
  • Verifying end to end data completeness, integrity, and accuracy during drill exercises, documenting results and remediation recommendations.
  • Identifying exit readiness gaps across infrastructure, deployment automation, monitoring, and data recovery processes, and driving corrective actions with engineering teams.
  • Designing, implementing, and supporting automated IaaC workflows utilizing Terraform, AWS CloudFormation, and standardized CI/CD pipelines.
  • Managing and optimizing Kubernetes clusters and containerized workloads (Docker), including cluster provisioning, scaling, and workload reliability improvements.
  • Building and maintaining observability solutions using CloudWatch, Datadog, and other monitoring/alerting tools to ensure service reliability and proactive incident response.
  • Developing automation, tooling, and scripts using Python or Java to reduce manual processes and enhance operational repeatability.
  • Collaborating with platform engineering, security, applications, and data teams to ensure consistent, secure, and compliant platform operations.
  • Participating in on‑call rotations, root cause analyses, and incident response activities to improve system resilience and operational excellence.
Qualifications

Required Qualifications

  • Bachelor’s degree and 8–10 years of relevant SRE, DevOps, cloud engineering, or infrastructure engineering experience; or 12 years of experience with a high school diploma. 
  • Expert level hands on knowledge of AWS services across compute, networking, storage, IAM, and serverless components.
  • Strong experience with Infrastructure as Code (Terraform, CloudFormation) and infrastructure automation principles.
  • Experience building CI/CD deployment pipelines and progressive delivery mechanisms using GitHub actions or similar tools
  •  Deep understanding of Kubernetes administration, container orchestration, and Docker based deployments.
  • Proven experience validating DR processes, performing system rebuilds, and conducting data integrity checks.
  • Experience building monitoring tools like dashboards, metrics, logs, and alerting systems using CloudWatch, Datadog, or similar observability tools.
  • Proficiency with programming/scripting languages such as Python, Java, or C# or Go.
  • Experience debugging complex failure modes, including cascading failures, network partitions, backpressure, and eventual consistency issues.
  • Strong analytical and documentation skills with the ability to clearly communicate technical findings to cross functional teams.
  • Ability to work in a fast-paced environment supporting high visibility, mission critical systems.
  • Ability to obtain a Public Trust clearance. 
  • US Citizen or Green Card Holder. 

Preferred Qualifications

  • AWS DevSecOps Engineer certification (preferred).
  • Additional AWS certifications (Solutions Architect, SysOps, Developer) and/or Kubernetes certifications (CKA, CKAD).
  • Familiarity with Zero Trust security models and cloud security best practices.
  • Experience with GitLab, Jenkins, or similar CI/CD platforms.
  • Experience with highly regulated environments (healthcare, finance, DHS, DoD, CMS, etc.).
  • Experience supporting federal, defense, or largescale enterprise programs involving legacy-to-cloud modernization.
  • Prior involvement in large‑scale DR drills, continuity of operations (COOP), or portability/executable readiness assessments.
Target Salary Range$112,000 - $179,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual’s experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.

Similar Jobs

More Jobs at Peraton

More Information Technology Jobs

Find similar Lead Site Reliability Engineer jobs: