ECS

Cloud Site Reliability Engineer (SRE)

ECS$130K — $180K *
US-AnywhereRemote in Arlington, VA
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent practical experience)
  • 5+ years of SRE experience (or equivalent), with demonstrated technical leadership
  • 10 years of general work experience
  • Experience building self-healing/auto-remediating systems
  • Expert knowledge of AWS, with GovCloud experience preferred
  • Deep expertise in Kubernetes and Terraform at production scale
  • Strong background in software engineering (Python and/or Go)
  • Experience with observability platforms (e.g., Grafana, Splunk, Prometheus, Loki)
  • Proven experience in incident command and postmortem reviews
  • Strong communication skills with both technical and federal leadership audiences
  • Ability to obtain/maintain required government clearance or suitability (CAC/PIV as applicable)
  • US Citizenship

Responsibilities

  • Design and build automated systems for self-healing operations
  • Shift operational focus to proactive automation rather than reactive fixes
  • Automate service restoration processes prior to root cause investigation
  • Define data-driven SLOs and error budgets in collaboration with service teams
  • Leverage live metrics to inform service reliability and investment priorities
  • Enforce infrastructure-as-code practices for all tech changes
  • Own Terraform standards and oversee reusable modules across programs
  • Implement a containerization-first approach using production-scale Kubernetes
  • Establish CI/CD and pipeline-as-code standards with progressive delivery
  • Build robust monitoring and logging systems for automation feedback
  • Manage incident response framework including escalation and post-incident analysis
  • Collaborate with development teams to integrate reliability and automation
  • Mentor engineers in adopting an automation-first culture
  • Ensure compliance with ATO/RMF and FedRAMP High standards for infrastructure

Benefits

  • Remote work flexibility available
  • Support for professional development and training
  • Health, dental, and vision insurance
  • Retirement savings plans
  • Paid time off and holidays
  • Relocation assistance may be available
  • Flexible work hours
Full Job Description
Everforth ECS is seeking a Cloud Site Reliability Engineer (SRE) to work in our Arlington, VA office/remotely.

About the Role
This role owns reliability and operational readiness for production systems across our federal cloud platform (AWS GovCloud, IL5 zero-trust). You'll define what "reliable enough" looks like for our services, build the automation that gets us there, and do it all on an infrastructure-as-code (IaC) foundation.

Responsibilities
Self-Healing Operations
  • Design and build automated remediation so systems detect, respond to, and recover from failure without manual intervention
  • Shift the team's posture from "is it running, how do we fix it" to "how do we make it fix itself"
  • Automate service restoration first; investigate root cause after

Uptime Goals & Reliability
  • Define reasonable, data-driven SLOs and error budgets for critical services alongside the teams that own them
  • Use live metrics to decide what's "reliable enough" and where to invest next

Infrastructure
  • Enforce infrastructure-as-code and configuration-as-code, no manual tech change
  • Own Terraform standards and reusable modules adopted across programs
  • Drive a containerization-first approach with production-scale Kubernetes (multi-tenancy, security policies, advanced scheduling)
  • Set CI/CD and pipeline-as-code standards, including progressive delivery

Observability & Incidents
  • Build monitoring, logging, alerting, and tracing (Datadog, Splunk) that gives automation the signal it needs to self-correct
  • Own the incident framework: escalation, restoration, root cause analysis, and post-incident review that closes the loop with more automation

Collaboration & Leadership
  • Partner with development and contractor teams leads to embed reliability and automation across the software
  • Mentor engineers toward this same automation-first philosophy
  • Support ATO/RMF and FedRAMP High compliance as it relates to infrastructure and automation


Salary Range: $130,000 - $180,000

General Description of Benefit

  • Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent practical experience)
  • 5+ years of SRE experience (or equivalent), with demonstrated technical leadership
  • 10 years of general work experience
  • Track record building self-healing/auto-remediating systems, not just dashboards
  • Jenkins experience
  • Expert AWS knowledge, GovCloud experience strongly preferred
  • Deep Kubernetes and Terraform expertise at production scale
  • Strong software engineering background (Python and/or Go)
  • Experience operating observability platforms (Grafana, Splunk, Prometheus, Loki, etc.)
  • Proven incident command and postmortem experience
  • Strong communication skills across technical and federal leadership audiences
  • Ability to obtain/maintain required government clearance or suitability (CAC/PIV as applicable)
  • US Citizenship

About ECS

ECS is a leading provider of digital solutions and services to the federal government. The company was founded in 2001 by Roy Kapani and has since grown to become a trusted partner to a wide range of government agencies. ECS offers a broad range of services, including cloud computing, cybersecurity, and artificial intelligence. The company has been recognized for its innovative solutions and has won numerous awards, including the AWS Public Sector Partner of the Year award.
Learn more about ECS
Size
2,000 employees
Industry

Similar Jobs

More Jobs at ECS

More Information Technology Jobs

Find similar Cloud Site Reliability Engineer (SRE) jobs: