Site Reliability Engineer

Prolaio

• $118K *
Healthcare
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in a technical field or equivalent experience.
  • 2-5 years in reliability engineering or related roles.
  • Proficient in SQL and Python for data analysis.
  • Experience with cloud data platforms, preferably Google Cloud and BigQuery.
  • Ability to clearly document technical investigations.

Responsibilities

  • Build and operate monitoring for data freshness and quality.
  • Define SLIs, SLOs, and maintain error budgets for data services.
  • Participate in on-call rotation and manage incident responses.
  • Investigate device failures end to end, correlating data effectively.
  • Analyze fleet behavior by clustering failures to identify trends.
  • Improve diagnostics processes by closing data analysis loops.
  • Develop reliability tooling using Python and SQL for automation.

Benefits

  • Competitive salary and performance bonuses.
  • Comprehensive health coverage including medical, dental, and vision.
  • Flexible spending options, including HSA and FSA.
  • Generous paid time off and holiday leave.
  • Paid parental and caregiver leave.
  • Company-paid life insurance and disability coverage.
  • 401(k) plan for long-term financial security.
  • Access to telehealth and supplemental coverage options.
Full Job Description
What Will You Do?

The Overview

The Site Reliability Engineer will ensure that Prolaio's cardiovascular data platform reliably captures, transports, and delivers continuous biosensor data from participants to the clinical researchers and trial sponsors who depend on it. In this role, you will build and operate the monitoring, service level objectives, incident response, and automation needed to keep critical data flows healthy, identify failures before they impact studies, and ensure that device downtime never goes unnoticed simply because a shipment record says a device was delivered.

This role is ideal for a self-starter who thrives at the intersection of software reliability, connected devices, and clinical data. You will investigate incidents across the full participant-to-platform stack, combining service telemetry with device data and participant reports to determine whether an issue originated with the hardware, phone, firmware, mobile application, connectivity, or the platform itself. You will turn those investigations into durable improvements, automating manual checks, strengthening observability, reducing operational toil, and building the reliability practices that keep Prolaio's clinical data complete and trustworthy at scale.

The Specifics
  • Build monitoring that treats missing data as a failure rather than a quiet day - freshness, volume and quality checks that verify delivery from outside the system instead of trusting a job's exit code.
  • Define and maintain SLIs, SLOs and error budgets for customer-facing data services, and report attainment honestly, including when the news is bad.
  • Join the on-call rotation, run incident response, and write blameless postmortems that produce owned follow-up actions.
  • Investigate device failures end to end: correlate discharge curves, Bluetooth disconnects and wear detection against the participant's complaint, and establish or exonerate each candidate cause on evidence.
  • Treat the fleet as a population - failure rates per device-month, clustering by build lot, firmware version or site - rather than as a queue of individual tickets.
  • Close the loop. When an analysis cannot reach a verdict, name the missing metric or threshold and route it back into the diagnostics pipeline so the next case is answerable from telemetry alone.
  • Build reliability tooling in Python and SQL on BigQuery, and keep monitors and probes defined as code rather than configured by hand.
  • Retire manual checks by automating them. Operational toil should shrink quarter over quarter, not accumulate.

Who You Are?
  • Bachelor's degree in Electrical Engineering, Computer Engineering, Computer Science, Biomedical Engineering or a related technical field, or equivalent practical experience.
  • 2-5 years in reliability engineering, production operations, hardware or systems test, field failure analysis, or data engineering - in a role where you were accountable for something that had to keep working.
  • Working SQL and Python: able to pull a dataset, characterize it, and defend the query behind a number.
  • Experience with a cloud data platform (Google Cloud and BigQuery preferred; AWS or Azure equivalents are fine).
  • Ability to write up a technical investigation clearly - what was observed, what was ruled out and how, and what remains unknown.

Additional Qualifications (Nice to Haves)
  • Hands-on exposure to Bluetooth Low Energy (GATT services and characteristics, connection intervals, pairing and bonding, disconnect behaviour), Wi-Fi or cellular data standards - and how each one fails in the field rather than how it is specified to behave.
  • Battery and power behaviour: discharge characterisation, charge-state interpretation, and telling a measurement artefact from a real capacity defect.
  • Formal reliability method - FMEA and risk priority numbers, FRACAS practice, and standards such as IEC 60812 or ISO 13485.
  • Observability tooling (Datadog, Grafana or similar) and synthetic monitoring.
  • Android device telemetry, mobile application diagnostics, or embedded firmware experience.
  • Prior work in medtech, digital health, clinical trials or another regulated, patient-facing environment.

Why You'll Love Working Here
  • Meaningful Compensation: Competitive salary, performance bonus, and equity so you can share in what we build.
  • Great Health Coverage: Medical, dental, and vision plans with multiple options and strong company contributions.
  • Flexible Spending Perks: HSA, FSA, commuter benefits, and a $1,200 annual Lifestyle Spending Account to support wellness, commuting, family needs, and more.
  • Time to Recharge: Generous paid time off, sick leave, and company holidays.
  • Family-First Benefits: Paid parental leave, caregiver leave, and support for growing families.
  • Security & Peace of Mind: Company-paid life insurance and short- and long-term disability coverage.
  • Plan for the Future: 401(k) plan to help you build long-term financial security.
  • Care When You Need It: Easy access to telehealth and optional supplemental coverage for life's unexpected moments.

Starting Salary is at $ 118,000.00 (Exact Compensation may vary based on skills, experience, and location)

Similar Jobs

More Jobs at Prolaio

More Healthcare Jobs

Find similar Site Reliability Engineer jobs: