Everforth ECS is seeking a
Cloud Site Reliability Engineer (SRE) to work in our
Arlington, VA office/remotely. About the Role This role owns reliability and operational readiness for production systems across our federal cloud platform (AWS GovCloud, IL5 zero-trust). You'll define what "reliable enough" looks like for our services, build the automation that gets us there, and do it all on an infrastructure-as-code (IaC) foundation.
Responsibilities Self-Healing Operations
- Design and build automated remediation so systems detect, respond to, and recover from failure without manual intervention
- Shift the team's posture from "is it running, how do we fix it" to "how do we make it fix itself"
- Automate service restoration first; investigate root cause after
Uptime Goals & Reliability
- Define reasonable, data-driven SLOs and error budgets for critical services alongside the teams that own them
- Use live metrics to decide what's "reliable enough" and where to invest next
Infrastructure
- Enforce infrastructure-as-code and configuration-as-code, no manual tech change
- Own Terraform standards and reusable modules adopted across programs
- Drive a containerization-first approach with production-scale Kubernetes (multi-tenancy, security policies, advanced scheduling)
- Set CI/CD and pipeline-as-code standards, including progressive delivery
Observability & Incidents
- Build monitoring, logging, alerting, and tracing (Datadog, Splunk) that gives automation the signal it needs to self-correct
- Own the incident framework: escalation, restoration, root cause analysis, and post-incident review that closes the loop with more automation
Collaboration & Leadership
- Partner with development and contractor teams leads to embed reliability and automation across the software
- Mentor engineers toward this same automation-first philosophy
- Support ATO/RMF and FedRAMP High compliance as it relates to infrastructure and automation
Salary Range: $130,000 - $180,000
General Description of Benefit
- Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent practical experience)
- 5+ years of SRE experience (or equivalent), with demonstrated technical leadership
- 10 years of general work experience
- Track record building self-healing/auto-remediating systems, not just dashboards
- Jenkins experience
- Expert AWS knowledge, GovCloud experience strongly preferred
- Deep Kubernetes and Terraform expertise at production scale
- Strong software engineering background (Python and/or Go)
- Experience operating observability platforms (Grafana, Splunk, Prometheus, Loki, etc.)
- Proven incident command and postmortem experience
- Strong communication skills across technical and federal leadership audiences
- Ability to obtain/maintain required government clearance or suitability (CAC/PIV as applicable)
- US Citizenship