Site Reliability Engineer

Seek Now

$110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4+ years in SRE, DevOps, or cloud operations roles supporting production systems
  • Deep experience with AWS services like EC2, ECS/EKS, Lambda
  • Proven incident management skills including on-call ownership and root cause analysis
  • Track record in building automated CI/CD pipelines using tools like GitHub Actions or Jenkins
  • Strong infrastructure-as-code capabilities with Terraform or CloudFormation
  • Proficient in at least one programming language such as Python or Go
  • Experience with container technologies like Docker and Kubernetes
  • Familiarity with observability tools like Datadog or Grafana

Responsibilities

  • Own the reliability and performance of production AWS services
  • Design and maintain automated CI/CD pipelines for seamless code deployment
  • Lead incident management efforts including outage coordination and root cause investigations
  • Define and monitor SLOs, SLIs, and error budgets with product and engineering teams
  • Enhance observability through effective monitoring and alerting strategies
  • Manage infrastructure using code to reduce manual work
  • Collaborate with development teams to incorporate reliability practices into systems

Benefits

  • Comprehensive health, dental, and vision coverage
  • 401(k) retirement plan with company match
  • Flexible PTO policy allowing for work-life balance
  • Hybrid work arrangement available in Atlanta
Full Job Description
The Role

As a Site Reliability Engineer, you'll be responsible for the availability, scalability, and operational health of our AWS-hosted infrastructure. You'll lead incident response, build the automation that lets our engineering teams ship safely and often, and drive a culture of measurable reliability across the organization.

What You'll Do
  • Own the reliability and performance of production services running in AWS, including capacity planning, cost optimization, and architecture reviews
  • Design, build, and maintain fully automated CI/CD pipelines that take code from commit to production with minimal manual intervention
  • Lead incident management: serve in the on-call rotation, coordinate response during outages, run blameless postmortems, and drive remediation to completion
  • Define and track SLOs, SLIs, and error budgets in partnership with product and engineering teams
  • Build and improve observability through monitoring, logging, alerting, and distributed tracing
  • Manage infrastructure as code and eliminate toil through automation
  • Partner with development teams to embed reliability best practices into system design and release processes
  • Contribute to disaster recovery planning, security hardening, and compliance efforts

What We're Looking For
  • 4+ years in SRE, DevOps, or cloud operations roles supporting production systems
  • Deep hands-on experience operating workloads in AWS (e.g., EC2, ECS/EKS, Lambda, RDS, S3, IAM, VPC, CloudWatch)
  • Proven experience with incident management: on-call ownership, incident command, root cause analysis, and postmortem processes
  • Demonstrated track record of building fully automated CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins, CodePipeline, or similar)
  • Strong infrastructure-as-code skills with Terraform, CloudFormation, or CDK
  • Proficiency in at least one scripting or programming language (Python, Go, Bash)
  • Experience with containers and orchestration (Docker, Kubernetes)
  • Familiarity with observability tooling such as Datadog, Prometheus/Grafana, or the ELK stack
  • Clear communicator who stays calm under pressure and can explain complex issues to technical and non-technical audiences

Nice to Have
  • AWS certifications (Solutions Architect, DevOps Engineer)
  • Experience with GitOps workflows (ArgoCD, Flux)
  • Background in security operations, compliance frameworks (SOC 2, ISO 27001)

What We Offer
  • Competitive salary
  • Comprehensive health, dental, and vision coverage
  • 401(k) with company match
  • Flexible PTO and hybrid work arrangement in Atlanta

Location

This role is based in Atlanta, GA, with a hybrid schedule.

Similar Jobs

More Jobs at Seek Now

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: