Staff Site Reliability Engineer

Hippocratic AI

$150K — $180K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years of experience in site reliability and software engineering
  • Bachelor's degree in Computer Science from a top program
  • Proficient in Python and/or Go for orchestration and scheduling
  • Experience designing systems that analyze operational metrics
  • Deep understanding of CI/CD and infrastructure automation
  • Hands-on experience with major cloud platforms (AWS, GCP, Azure)
  • Strong knowledge of containerization tools like Docker and Kubernetes

Responsibilities

  • Design and build GPU management and scheduling platform for inference
  • Develop metrics pipeline for GPU load and utilization data collection
  • Implement admission control for managing inference requests
  • Create autoscaling mechanisms for model replicas based on demand
  • Architect scalable, fault-tolerant production systems in cloud environments
  • Build infrastructure automation and deployment pipelines using Terraform
  • Maintain monitoring and logging for platform reliability and performance
  • Enforce security and compliance policies in healthcare AI contexts
  • Collaborate with engineers to resolve complex operational issues
  • Mentor junior engineers and improve team technical skills

Benefits

  • Opportunity to work on high-impact projects in healthcare AI
  • Mentorship and professional development opportunities
  • Collaborative work environment with engineers and research scientists
  • Ownership of complex systems and engineering challenges
  • Flexible work arrangements to boost work-life balance
  • Engagement with cutting-edge AI technologies and practices
Full Job Description
About the Role

We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on - and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently managing a large fleet of GPU-backed models.

We run nearly 30 models across heterogeneous hardware, and keeping that fleet fast, reliable, and cost-effective is a serious engineering challenge. You'll build the GPU management and scheduling platform that sits at the center of it - collecting utilization and load metrics, interpreting what they actually mean, and using them to make real-time decisions about admission control and scaling. The goal: route and schedule inference calls so we use our capacity efficiently without exceeding it, and scale model replicas up and down automatically as demand shifts.

This is a senior role for someone with a decade in the field who can move fluidly between systems engineering and software development, and who is excited to own a complex, evolving system end to end.

What You'll Do
  • Design and build our GPU management and scheduling platform - the system that decides when, where, and how inference calls run across a fleet of ~30 models on heterogeneous hardware
  • Build the metrics pipeline that collects GPU load and utilization data, and the logic that turns those signals into decisions
  • Implement admission control to protect capacity - deciding when to accept, queue, or shed inference requests so we operate within fleet limits
  • Build autoscaling that adjusts the number of model replicas in response to real-time demand and utilization
  • Develop cloud orchestration systems and operators in Python and Go to manage the model fleet
  • Architect and operate scalable, fault-tolerant, secure production systems on AWS, GCP, or Azure
  • Design and build infrastructure automation and deployment pipelines (Terraform, CI/CD) as first-class software
  • Stand up and maintain monitoring, logging, and alerting that keep the platform reliable and performant
  • Develop and enforce security and compliance policies appropriate to a healthcare AI platform
  • Partner with engineers and research scientists to diagnose and resolve complex infrastructure, deployment, and operational issues
  • Mentor engineers and raise the technical bar across the team
What You Bring

Must-Have
  • 10+ years of professional experience across site reliability / DevOps engineering and software engineering
  • Computer Science Degree Required from a top CS program.
  • Strong software engineering fundamentals - you build orchestration and scheduling systems in Python and/or Go, not just configure off-the-shelf tools
  • Experience designing systems that make decisions from operational metrics - collecting signals, interpreting them, and driving control loops such as autoscaling, load shedding, or admission control
  • Deep experience with infrastructure automation and CI/CD (Terraform, GitLab CI/CD, or similar)
  • Hands-on production experience with at least one major cloud platform (AWS, GCP, or Azure)
  • Strong knowledge of containerization and orchestration (Docker, Kubernetes)
  • Experience with monitoring and logging stacks (ELK, Grafana, Datadog, or similar)
  • Familiarity with secrets management and security tooling (HashiCorp Vault, AWS KMS, Azure Key Vault)
  • Excellent problem-solving skills and the ability to work both independently and collaboratively
  • Strong communication and interpersonal skills

Nice-to-Have
  • Experience managing GPU fleets or scheduling workloads across heterogeneous accelerators
  • Familiarity with ML inference serving and model deployment (e.g. Triton, KServe, Ray Serve, or similar)
  • Experience with Kubernetes autoscaling internals (HPA/VPA, custom metrics, custom controllers)
  • Experience implementing HIPAA and SOC 2 compliance
  • Experience operating in an HPC environment
  • Bachelor's or Master's in Computer Science, Computer Engineering, or a related field


Join our team at Hippocratic AI and help shape the future of clinically safe, production-grade AI systems.

Similar Jobs

More Jobs at Hippocratic AI

More Information Technology Jobs

Find similar Staff Site Reliability Engineer jobs: