GovCIO

Site Reliability Engineer

GovCIO$230K — $250K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree with 12+ years in cloud/infrastructure engineering (or equivalent experience)
  • 5-10+ years engineering experience focusing on Linux and Windows systems
  • Expertise in Kubernetes and container-based platforms
  • Experience with cloud infrastructure environments
  • Proficient in scripting languages like Python and Go
  • Hands-on experience with Terraform and other automation tools
  • Familiarity with monitoring platforms and incident management

Responsibilities

  • Design and maintain highly available production systems
  • Define and manage SLIs, SLOs, and error budgets
  • Automate operational tasks to reduce manual processes
  • Develop solutions for monitoring, alerting, and observability
  • Enhance system performance, capacity, and resilience
  • Lead incident response efforts and conduct root cause analysis
  • Implement disaster recovery and business continuity strategies
  • Collaborate with development teams to boost application reliability

Benefits

  • Hybrid/remote working options
  • Opportunity to work on mission-critical systems
  • Focus on reliability and scalability in engineering efforts
  • Potential for career growth in a dynamic environment
  • Collaborative culture with an emphasis on DevOps practices
Full Job Description
Overview

GovCIO is hiring a Site Reliability Engineer with an active Secret clearance to ensure reliability, scalability, performance, and availability of mission-critical systems by combining software engineering practices with infrastructure operations expertise. This role is based in Arlington, VA, as a hybrid/remote position.

Responsibilities

Responsibilities:

  • Design and maintain highly available production systems.
  • Define and manage SLIs, SLOs, and error budgets.
  • Automate operational tasks and eliminate manual processes.
  • Develop monitoring, alerting, and observability solutions.
  • Improve system performance, capacity, and resilience.
  • Lead incident response and root cause analysis.
  • Implement disaster recovery and continuity strategies.
  • Partner with development teams to improve application reliability.
Qualifications

Qualifications:

Bachelor's with 12+ years of infrastructure/cloud engineering experience (or commensurate experience)

Required Skills and Experience:

  • 510+ years of engineering experience, with a strong background in Linux and Windows systems
  • Expertise in Kubernetes and container platforms
  • Experience working with cloud infrastructure environments
  • Proficiency in scripting languages such as Python and Go
  • Hands-on knowledge of Terraform and automation tools
  • Familiarity with monitoring platforms and incident management practices
  • Experience designing and managing CI/CD pipelines

Preferred Skills and Experience:

  • Kubernetes certifications
  • AWS/Azure certifications
  • DevOps certifications
  • ITIL preferred

Clearance Required: Must have an active Secret clearance and be able to attain DEA suitability.

Posted Salary RangeUSD $230,000.00 - USD $250,000.00 /Yr.

About GovCIO

GovCIO is a technology and consulting firm that provides IT solutions to government agencies. The company specializes in cloud computing, cybersecurity, and digital transformation. GovCIO's mission is to help government agencies improve their IT infrastructure and enhance their services to the public. The company was founded in 2015 and is headquartered in Washington, DC.
Learn more about GovCIO
Size
50 employees
Industry
Founded
2015

Similar Jobs

More Jobs at GovCIO

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: