Reliability Engineer (onsite)

System One Holdings, LLC

$130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4+ years in cloud infrastructure, systems engineering, DevOps, SRE, or similar roles
  • 2+ years of hands-on Microsoft Azure experience
  • 2+ years of hands-on Terraform expertise for production infrastructure
  • Experience administering/supporting Kubernetes (preferably AKS)
  • Knowledge of cloud networking concepts and troubleshooting techniques
  • Strong analytical and problem-solving abilities
  • Ability to work on-site in Atlanta, GA or Washington DC area

Responsibilities

  • Design, implement, maintain, and troubleshoot production Azure infrastructure using Terraform
  • Support reliability, performance, and availability of workloads in Azure Kubernetes Service (AKS)
  • Troubleshoot cloud infrastructure, networking, Kubernetes, and application reliability issues
  • Automate cloud operations to reduce manual work and improve consistency
  • Collaborate with development and operations teams to enhance deployment and incident response practices
  • Implement and refine monitoring, alerting, dashboards, and operational reporting
  • Document infrastructure, procedures, troubleshooting guidance, and operational runbooks
  • Identify reliability risks and recommend improvements to cloud architecture and processes

Benefits

  • Health and welfare benefits coverage options including medical, dental, and vision
  • Spending accounts and life insurance options
  • Participation in a 401(k) plan
  • Voluntary plans available
Full Job Description
Reliability Engineer
Washington, DC or Atlanta, GA - ONSITE
$130,000.00
Must be U.S. Citizen or Lawful Permanent Resident (Green Card Holder) per government contract
Security Clearance: Public Trust


In this role, the engineer will ensure the reliability, scalability, and operational health of EDAV's Azure cloud environment. Terraform is central to this position: the engineer will independently design, build, review, and troubleshoot Infrastructure-as-Code for mission critical environments. The role involves close collaboration with platform engineers, developers, security teams, and product stakeholders to automate cloud infrastructure, improve Kubernetes operations, and resolve issues impacting the availability of EDAV data and analytics services.

Job Responsibilities:
• Design, implement, maintain, and troubleshoot production Azure infrastructure using Terraform.
• Support reliability, performance, and availability of workloads in Azure Kubernetes Service (AKS).
• Troubleshoot cloud infrastructure, networking, Kubernetes, and application reliability issues.
• Automate cloud operations to reduce manual work and improve consistency.
• Collaborate with development and operations teams to enhance deployment and incident response practices.
• Implement and refine monitoring, alerting, dashboards, and operational reporting.
• Identify reliability risks and recommend improvements to cloud architecture and processes.
• Document infrastructure, procedures, troubleshooting guidance, and operational runbooks.

Job Requirements:
• 4+ years in cloud infrastructure, systems engineering, DevOps, SRE, or similar roles
• 2+ years of hands on Microsoft Azure experience
• 2+ years of hands on Terraform expertise to design reusable modules, manage state, troubleshoot failures, and maintain production infrastructure
• Experience administering/supporting Kubernetes (preferably AKS)
• Experience supporting cloud hosted systems and automating cloud operations
• Knowledge of cloud networking concepts and troubleshooting
• Possession of strong analytical and problem-solving abilities
• Ability to work on-site in Atlanta, GA or the Washington DC metro-area
• Ability to obtain/maintain Public Trust/Suitability clearance
• Bachelor's degree or equivalent experience
• Must be U.S. Citizen or Lawful Permanent Resident (Green Card Holder)

Job Desirables:
• Experience with certificate lifecycle management (maintaining, renewing, rotating certs)
• Experience creating operational dashboards or reports (Power BI preferred)
• Experience with observability platforms (Grafana, Prometheus, Elastic, Splunk)
• Experience integrating Azure resources with Active Directory
• Experience using agentic coding tools (Claude Code, Codex, GitHub Copilot) for applications or infrastructure automation
• Azure, Kubernetes, or Terraform certifications

System One not only serves as a valued partner for our clients, but we offer eligible employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.

#M-1
#LI-VH1

Ref: #851-Rockville-S1

Similar Jobs

More Jobs at System One Holdings, LLC

More Information Technology Jobs

Find similar Reliability Engineer (onsite) jobs: