Site Reliability Manager

Karsun Solutions, LLC

$155K — $175K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field; Master's degree preferred
  • 10+ years in a similar role managing a site reliability team in the AWS cloud
  • 5+ years supporting operations for fault-tolerant, cloud-native production applications
  • Deep understanding of AWS and containerization technologies (e.g., Docker, Kubernetes)
  • Strong knowledge of infrastructure as code tools (e.g., Terraform, Ansible) and CI/CD practices
  • Experience with Datadog for monitoring and observability
  • Excellent communication skills and ability to collaborate with cross-functional teams
  • Ability to obtain and maintain a Public Trust clearance.

Responsibilities

  • Lead and mentor a 8-20 person service delivery team in application reliability and DevSecOps.
  • Ensure production reliability and standardize observability using Datadog and develop key performance metrics.
  • Implement best practices for infrastructure as code and deployment automation within DevSecOps.
  • Oversee platform lifecycle, collaborating on scalable, fault-tolerant architecture design.
  • Drive continuous improvement initiatives to enhance system efficiency.

Benefits

  • Opportunity for professional development and mentoring
  • Work in a dynamic environment focusing on learning and innovation
  • Engagement with cutting-edge technologies in cloud services
  • Collaborative team culture with cross-functional projects
  • Eligibility for a public trust security clearance.
Full Job Description
Summary

We are seeking a highly skilled and experienced Site Reliability Manager to join our team to ensure the reliability, scalability, and performance of our systems and services. You will lead a team of engineers focusing on three core pillars: Application Reliability, DevSecOps, and Platform Lifecycle Management. The ideal candidate must reside in DMV area and be available to work on site in office or customer locations in this area. Must have demonstrated experience in having performed this role for at least 3 years.

What You'll Be Doing:
  • Team Leadership: Lead an 8-20 person service delivery team (Service Support specialist, DevSecOps, and Site Reliability engineers), mentoring them to foster a culture of learning and innovation.
  • Pillar 1: Application Reliability (Core Focus): Take joint ownership of production reliability, standardize observability and error handling using Datadog, develop SLOs and KPIs to measure system performance, and conduct incident post-mortems and root cause analyses.
  • Pillar 2: DevSecOps: Define and implement best practices for infrastructure as code, deployment automation, and drive vulnerability management and resolution.
  • Pillar 3: Platform Lifecycle Management: Oversee the end-to-end platform lifecycle, collaborate with cross-functional teams to design scalable and fault-tolerant architectures, and drive continuous improvement initiatives to enhance system efficiency.

Required Qualifications:
  • Bachelor's degree in Computer Science, Engineering, or a related field; Master's degree preferred.
  • 10+ years of experience in a similar role managing a team of site reliability engineers and delivering in the AWS cloud platform.
  • 5+ years of experience supporting operations and maintenance for cloud-native applications in production that are fault-tolerant, self-healing, scalable, and highly available.
  • Deep understanding of the AWS cloud computing platform and containerization technologies (e.g., Docker, Kubernetes).
  • Strong knowledge of infrastructure as code tools (e.g., Terraform, Ansible, ArgoCD) and CI/CD pipelines.
  • Experience with Datadog as the primary logging, monitoring, and observability platform (alongside AWS Cloudwatch).
  • Excellent communication and interpersonal skills, with the ability to collaborate effectively with cross-functional teams.
  • Strong problem-solving and analytical skills, with a keen attention to detail.
  • Ability to obtain and maintain a Public Trust clearance.
  • Certifications such as AWS Certified DevOps Engineer are a plus.

Preferred Qualifications:
  • Understanding of modern architecture, e.g., micro-services, EDA, etc., and a cautious approach against overcomplexity and overengineering.
  • Experience designing and operating distributed systems and cloud infrastructure at scale.


Salary Range

The proposed salary range for this role is $155,000 to $175,000 USD. The salary range provided is a good faith estimate representative of all experience levels. Karsun considers several factors when extending an offer, including but not limited to, the role, function and associated responsibilities, a candidate's work experience, location, education/training, and key skills.

Third Party Resumes

Karsun does not accept unsolicited resumes through or from search firms or staffing agencies. All unsolicited resumes will be considered the property of Karsun and Karsun will not be obligated to pay a placement fee.

Clearance Information

This position requires the eligibility to obtain a security clearance. The Defense Industrial Security Clearance Office (DISCO), an agency of the Department of Defense, handles and adjudicates the security clearance process. More information about Security Clearances can be found on the US Department of State government website: https://www.state.gov/m/ds/clearances/c10978.htm

Location

To be considered for this role, you must reside in one of the following states: CA, CO, DC, FL, GA, IL, MD, NJ, NY, NC, OH, OK, PA, SC, TX, VA, WV.

Work Authorization

Applicants must be authorized to work in the U.S. We may consider candidates currently in H-1B status who are eligible for transfer.

Statement on AI and Hiring Process

At Karsun, we are committed to a fair and equitable hiring process. We do not use artificial intelligence to make decisions on candidates, scan resumes, or source potential hires. Applicants are reviewed by our talent acquisition team and hiring managers to ensure a thorough evaluation based on skills, experience, and alignment with our company values. Our hiring decisions are made with human judgement, ensuring fairness and transparency throughout the process.

Similar Jobs

More Jobs at Karsun Solutions, LLC

More Information Technology Jobs

Find similar Site Reliability Manager jobs: