ECS

Senior Site Reliability Engineer

ECS$118K — $177K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years of experience as a Site Reliability Engineer (SRE) or equivalent
  • 6+ years of experience in designing, implementing, and maintaining observability solutions
  • 6+ years of hands-on experience with SRE tools like Elastic, Prometheus, Grafana, Splunk
  • 3+ years experience defining and measuring Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
  • Proficient in programming or scripting with languages like Python and Bash
  • Strong knowledge of microservices, Docker, and Kubernetes
  • Proven ability to collaborate across cross-functional teams

Responsibilities

  • Define and implement the SRE practice to enhance reliability and performance
  • Identify areas for system improvements and drive initiatives for enhancement
  • Design and maintain logging, monitoring, and alerting systems using the Elastic stack
  • Respond to incidents and perform root cause analysis for incidents
  • Collaborate with developers and engineers to integrate reliability into the development lifecycle
  • Set and measure SLOs and SLIs for solutions
  • Enhance system scalability and efficiency through proactive improvements

Benefits

  • Working remotely from anywhere in the U.S.
  • Opportunity to influence and grow the SRE practice within an organization
  • Collaboration with cross-functional teams to enhance service reliability
  • Engagement in initiatives focused on continuous improvement
Full Job Description
Everforth ECS is seeking a Senior Site Reliability Engineer to work remotely.

Role & Responsibilities:

ECS is seeking a talented Senior Site Reliability Engineer (SRE) to play a key role in defining, implementing, and growing our SRE practice to ensure the reliability, availability, and performance of our critical production environments.

The Senior SRE will contribute to a culture of continuous improvement, identifying areas for enhancement, and driving initiatives to improve system reliability, scalability, and efficiency.

The successful candidate will have demonstrated hands-on experience designing, implementing, and maintaining solutions to ensure that systems, including infrastructure and applications, are resilient, highly available, and performant. The Senior SRE will also play a critical role in defining and measuring the Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for our solution.

The Senior SRE will be responsible for setting up comprehensive logging, monitoring, and alerting solutions using the Elastic stack and other tools as necessary to ensure the continuous performance of services. Additionally, they will respond to incidents, perform root cause analyses, and implement solutions to prevent reoccurrences. The Senior SRE will work in close collaboration with other SRE team members, developers, testers, infrastructure engineers, DevOps engineers, and other stakeholders to integrate reliability and observability into the software development lifecycle.

Salary Range: $118,000 - $177,000

General Description of Benefits

  • Must be a US citizen with the ability to obtain Public Trust Suitability.
  • 6+ years of experience as a Site Reliability Engineer (SRE) or equivalent
  • 6+ years of demonstrated experience designing, implementing, and maintaining observability solutions to include logging, monitoring, and alerting
  • 6+ years of hands-on experience with SRE tools (e.g., Elastic, Prometheus, Grafana, Splunk, etc.)
  • 3+ years defining and measuring SLOs and SLIs
  • 3+ years of relevant experience using cloud platforms (AWS GovCloud preferred)
  • 3+ years of hands-on programming or scripting (e.g., Python, Bash, etc.)
  • Strong knowledge of microservices, containerization, and orchestration tools (Docker, Kubernetes)
  • Proven ability to collaborate with cross-functional teams (development, testing, and product) to integrate reliability and observability into the software development lifecycle
  • Strong problem-solving and analytical skills
  • Proactive, detail-oriented approach to identifying inefficiencies and implementing improvements.
  • Proficient in developing Synthetic monitoring scripts using typescript.

About ECS

ECS is a leading provider of digital solutions and services to the federal government. The company was founded in 2001 by Roy Kapani and has since grown to become a trusted partner to a wide range of government agencies. ECS offers a broad range of services, including cloud computing, cybersecurity, and artificial intelligence. The company has been recognized for its innovative solutions and has won numerous awards, including the AWS Public Sector Partner of the Year award.
Learn more about ECS
Size
2,000 employees
Industry

Similar Jobs

More Jobs at ECS

More Enterprise Technology Jobs

Find similar Senior Site Reliability Engineer jobs: