Staff Site Reliability Engineer

Pivotal Health

$135K — $160K *
US-AnywhereRemote in Brooklyn, NY
Healthcare
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years of experience in site reliability engineering or large-scale production systems
  • Deep knowledge of cloud infrastructure, distributed systems, and modern deployment practices
  • Experience in designing and operating highly available systems in a fast-paced environment
  • Strong skills in observability, incident management, and performance engineering
  • Hands-on experience in creating automation tools to reduce operational toil
  • Effective at influencing architecture practices across teams without formal authority
  • Ability to communicate operational risks and technical trade-offs with diverse stakeholders

Responsibilities

  • Set Pivotal's reliability strategy and define service-level objectives
  • Design resilient cloud infrastructure to support scalability and complex workflows
  • Build a cohesive observability strategy with metrics, logs, and alerting mechanisms
  • Improve incident response practices and lead resolution of complex incidents
  • Automate repetitive manual tasks to enhance operational efficiency
  • Embed reliability practices across engineering teams for independent service building
  • Strengthen security and compliance practices for sensitive data handling

Benefits

  • Competitive compensation with equity options
  • Comprehensive health, dental, and vision insurance
  • 401(k) retirement savings plan
  • Flexible time off policies
  • Company-wide connection opportunities and events
Full Job Description
About the Role

We're hiring a Staff Site Reliability Engineer to define and strengthen how reliability, scalability, and operational excellence are built into Pivotal's platform.

This is a senior individual contributor role with broad influence across engineering. You'll work hands-on with our infrastructure and production systems while setting the technical direction for reliability across the company. You'll help us evolve from a rapidly growing platform into one that can scale predictably, recover gracefully, and meet the high standards of availability, security, and auditability required in healthcare.

You'll partner closely with software, data, AI, security, and product teams to design resilient systems, improve observability, reduce operational risk, and make reliability a shared engineering responsibility. You'll bring the technical depth to solve our most difficult infrastructure challenges and the leadership to establish practices that raise the bar for the entire organization.

If you're energized by complex distributed systems, greenfield infrastructure work, and the opportunity to shape reliability at a pivotal stage of company growth, this is the role for you.

What You'll Do
  • Set Pivotal's reliability strategy: Define the technical vision and roadmap for reliability, availability, scalability, and operational readiness. Establish clear service-level objectives and help teams make informed tradeoffs between reliability, velocity, and cost.
  • Design resilient, scalable infrastructure: Guide and implement improvements to our cloud architecture, deployment systems, networking, compute, storage, and other shared infrastructure. Ensure our systems can scale with increasing product usage, data volume, and workflow complexity.
  • Build world-class observability: Develop a cohesive approach to metrics, logs, traces, dashboards, and alerting. Give engineers the visibility they need to understand system behavior, identify emerging issues, and resolve production incidents quickly.
  • Improve incident response and resilience: Establish effective incident management, on-call, postmortem, and disaster recovery practices. Lead the response to complex incidents and ensure lessons result in durable improvements to our systems and processes.
  • Reduce operational toil through automation: Identify recurring manual work and build systems, tooling, and automation that make operating Pivotal's platform safer and more efficient. Improve deployment workflows, capacity management, infrastructure provisioning, and production diagnostics.
  • Embed reliability across engineering: Partner with software, data, and AI engineers to improve system design, production readiness, and failure handling. Create reusable patterns, tooling, and standards that allow teams to build reliable services without becoming dependent on a centralized operations function.
  • Strengthen security and compliance: Work closely with security and compliance stakeholders to protect sensitive healthcare and financial data. Help ensure our infrastructure, access controls, audit trails, and operational practices meet applicable regulatory and customer requirements.
  • Provide technical leadership: Serve as a trusted technical partner to senior engineers and engineering leaders. Lead architecture reviews, mentor engineers, clarify complex tradeoffs, and raise the standard for infrastructure and operational engineering across the organization.
Who You Are
  • 8+ years of experience in site reliability engineering, infrastructure engineering, platform engineering, or operating large-scale production systems
  • Deep knowledge of cloud infrastructure, distributed systems, networking, containers, orchestration, infrastructure as code, and modern deployment practices
  • Experience designing, building, and operating highly available systems in a fast-growing production environment
  • Strong understanding of observability, service-level objectives, capacity planning, incident management, disaster recovery, and performance engineering; Hands-on and technically credible, with the ability to debug complex issues across application, infrastructure, network, and data layers
  • Experienced in creating automation and internal tooling that reduces operational toil and improves developer productivity
  • Comfortable influencing architecture and engineering practices across teams without relying on formal authority
  • A thoughtful communicator who can translate operational risk and technical tradeoffs for engineering, product, security, and business stakeholders; Pragmatic about balancing reliability, delivery speed, complexity, and cost
  • Comfortable operating in ambiguity and building from scratch-you're energized by greenfield work, not slowed down by it


Extra credit if you have
  • Experience operating infrastructure that handles healthcare, financial, or other sensitive and regulated data
  • Familiarity with HIPAA, PHI, SOC 2, data privacy, auditability, and security controls
  • Experience supporting data-intensive, AI, or machine learning systems in production
  • Experience with multi-region architecture, high-volume asynchronous workflows, or complex distributed processing systems
  • A track record of introducing SRE practices at an early-stage or high-growth company


Benefits Include:
  • Competitive compensation, including equity
  • Full health, dental, and vision coverage
  • Retirement savings plan through 401(k)
  • Flexible time off
  • Opportunities for company-wide connection and events

Ready to Make an Impact?We're building something meaningful; and we want you on the team.

Bring your ideas, curiosity, and drive, and let's transform healthcare reimbursement together.

Employment Information

Work Authorization

Candidates must be authorized to work in the United States without current or future employer sponsorship.

Similar Jobs

More Jobs at Pivotal Health

More Healthcare Jobs

Find similar Staff Site Reliability Engineer jobs: