CentralReach

Sr. Site Reliability Engineer

CentralReach$160K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in Site Reliability Engineering or a related field.
  • Proficient in monitoring and observability tools like Splunk and Prometheus.
  • Strong grasp of CI/CD principles and tools such as Jenkins and GitHub Actions.
  • Deep understanding of AWS and cloud-native infrastructures.
  • Solid knowledge of containerization with Kubernetes and Helm.
  • Programming experience in languages like Java, Python, or Go, plus familiarity with .NET.
  • Understanding of Linux and Windows system environments.

Responsibilities

  • Own production reliability, including availability and performance of services.
  • Define and enhance reliability metrics such as SLOs and SLIs.
  • Troubleshoot and resolve operational issues impacting service reliability.
  • Automate observability capabilities and capacity forecasting.
  • Increase development velocity by reducing toil through automation.
  • Manage production support including incident and root cause analysis.
  • Collaborate with software engineering on release readiness and operational metrics.

Benefits

  • Comprehensive health benefits for full-time employees.
  • Generous paid time off (PTO) policy.
  • 401(k) matching program to support retirement savings.
  • Flexible hybrid work schedules.
  • Career development support for professional growth.
  • Wellness programs to promote employee well-being.
  • Opportunities for community engagement through CR Cares™ initiative.
Full Job Description
The Platform Engineering group at CentralReach builds the underlying technologies that power our Public and Private Cloud Platforms worldwide. The group is responsible for storage, data infrastructure, IT, observability systems, DevOps, SRE, provisioning, compute, orchestration platform, internal tools, internal platforms (laptops, networks, systems etc.) and services - all the components that make up the CentralReach Platform.

If you have a passion for the future, enjoy and thrive in an agile, fast-moving, ever-changing startup environment, welcome and take on technical challenges of all shapes and sizes, have excellent interpersonal skill and sense of humor and enjoy rolling up your sleeves and jumping in, then read on!

As a Sr. SRE, you will work closely with the key stakeholders in Software Engineering to drive adoption of modern reliability practices like SLOs, error budget policies, actionable alerts, incident retrospectives, chaos testing, and end-to-end ownership.

Key Accountabilities:
  • Own production reliability, including availability, latency, performance, capacity planning, monitoring, emergency response, and uptime for production environments.
  • Define, maintain, and improve SLOs, SLIs, error budgets, actionable dashboards, and observability practices.
  • Analyze, troubleshoot, and resolve operational issues that affect service reliability and SLO performance.
  • Build and automate multi-environment observability capabilities, including capacity forecasting based on usage patterns.
  • Reduce toil and increase development velocity through automation and continuous improvement.
  • Provide production support, including incident, change, and problem management; root cause analysis; service restoration; runbooks; and standard operating procedures.
  • Identify data-driven opportunities to improve system architecture, availability, performance, and reliability.
  • Collaborate with software engineering teams on release management, roadmap planning, and operational readiness.
  • Implement and manage reliability and observability tools such as Datadog, Prometheus, and Grafana.


Desired Skills and Experience:
  • Experience with monitoring, APM, and observability tools such as Splunk, Prometheus, Datadog, and OpenTelemetry.
  • Experience implementing observability strategies for logs, metrics, and traces.
  • Strong understanding of CI/CD practices and tools such as Jenkins, GitHub Actions, GitLab, Argo, and Kargo.
  • Strong understanding of major cloud providers, preferably AWS, and cloud-native infrastructure concepts.
  • Strong understanding of containerization technologies, including Kubernetes and Helm.
  • Experience with one or more programming languages, such as Java, Python, or Go, and familiarity with .NET application development.
  • Strong understanding of Linux, Windows, software development, systems, networking, and cloud concepts.
  • Experience using AI to improve productivity and amplify technical skills.


Base Salary Range

$160,000-$180,000 USD

We offer competitive compensation, comprehensive health benefits, generous PTO, 401(k) matching, and paid parental leave to our full-time employees. Our team members also enjoy hybrid work schedules, career development support, wellness programs, and opportunities to give back through CR Cares™, our community engagement initiative.

Be part of a market leader driving the future of care. Explore opportunities at centralreach.com/careers.

About CentralReach

CentralReach is a healthcare technology company that provides software and services to help healthcare providers improve patient outcomes. The company was founded in 2012 by Charlotte Fudge and Chris Sullens and is headquartered in Boca Raton, Florida. CentralReach's software platform includes tools for electronic health records, practice management, and data analysis, among others. The company serves a wide range of healthcare providers, including behavioral health clinics, speech and occupational therapy practices, and schools.
Learn more about CentralReach
Size
500 employees
Industry
Net Income
-$1 million
Founded
2012
5 Year Trend
+50%
Revenue
$20 million

Similar Jobs

More Jobs at CentralReach

More Information Technology Jobs

Find similar Sr. Site Reliability Engineer jobs: