Senior Site Reliability Engineer (Upmarket)

Heidi Health

$130K — $160K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3-6+ years in Site Reliability Engineering (SRE), DevOps, or operations-heavy engineering roles.
  • Experience in supporting production systems with on-call rotation participation.
  • Ability to debug live systems effectively under pressure.
  • Proficient in operating cloud infrastructure, particularly AWS.
  • Working knowledge of Kubernetes and containerized applications.
  • Experience with Infrastructure as Code tools like Terraform.
  • Familiarity with monitoring tools such as Datadog or Prometheus.

Responsibilities

  • Participate in on-call and incident response to manage production incidents and support effective communication.
  • Improve operational reliability by identifying issues and implementing fixes through automation and process changes.
  • Own and enhance Kubernetes clusters and cloud infrastructure as familiarity grows.
  • Strengthen system observability with improved dashboards and alert mechanisms.
  • Reduce operational toil by automating repetitive tasks and simplifying operational processes.
  • Support safe change through improved deployment processes and operational readiness.
  • Contribute to writing and maintaining runbooks while participating in blameless post-mortems.

Benefits

  • Collaborative in-office environment with like-minded professionals.
  • Healthcare, dental, and vision benefits.
  • 401k plan with a 3% company match.
  • Annual personal development budget of $500.
  • Equity options that allow employees to share in the company's success.
  • Opportunities to create a global impact in a leading healthtech startup.
  • Potential for accelerated career development within a startup environment.
Full Job Description
What you'll do
  • Participate in on-call and incident response:
    Respond to production incidents, contribute to service restoration, and support clear communication during incidents. Over time, take increasing responsibility for leading incidents end-to-end.
  • Improve operational reliability:
    Identify recurring issues and reliability risks, and drive fixes through better alerting, automation, system changes, or process improvements.
  • Own parts of the production environment:
    Operate and improve Kubernetes clusters, cloud infrastructure, and core platform services, with growing ownership as familiarity increases.
  • Strengthen observability:
    Improve dashboards, alerts, logs, and traces so issues are detected earlier and diagnosed faster, with a strong focus on actionable signals.
  • Reduce operational toil:
    Automate repetitive tasks, simplify runbooks, and improve tooling to make on-call and day-to-day operations easier and safer.
  • Support safe change:
    Improve deployments, rollback mechanisms, and operational readiness to reduce the risk of incidents caused by change.
  • Contribute to operational practices:
    Write and maintain runbooks, participate in blameless post-mortems, and help improve incident response processes over time.
  • Collaborate closely with engineers:
    Work with product and feature teams to improve production readiness, service ownership, and reliability expectations.


What we're looking for
  • 3-6+ years in SRE, DevOps, Platform, or operations-heavy engineering roles.
  • Experience supporting production systems and participating in on-call rotations.
  • Comfortable debugging live systems under pressure.
  • Experience operating cloud infrastructure (AWS preferred).
  • Working knowledge of Kubernetes and containerised workloads.
  • Infrastructure as Code experience (Terraform or similar).
  • Familiarity with monitoring and alerting tools (Datadog, Prometheus, etc).
  • Scripting or automation experience (Python, Bash, or similar).


The way we work

1. Build to Last

We design for safety and reliability so clinicians, patients, and our teams can trust what we build every day.

2. Own Your Practice

Ideas rise on merit, not title, and everyone shares responsibility for the standards we set together.

3. Move Fast, Stay Steady

We move quickly but never at the cost of trust. Progress only matters if people can depend on what we make.

4. Make Others Better

Honest feedback, steady support, and shared growth keep our teams improving together.

Why you will flourish with us ?
  • In office to collaborate with like-minded professionals
  • Healthcare, Dental, Vision benefit options
  • 401k with 3% match
  • Personal development budget of $500 per annum
  • Become an owner, with shares (equity) in the company, if Heidi wins, we all win
  • The rare chance to create a global impact as you immerse yourself in one of the leading healthtech startups
  • The opportunity to fast track your startup career!

Similar Jobs

More Jobs at Heidi Health

More Information Technology Jobs

Find similar Senior Site Reliability Engineer (Upmarket) jobs: