About the RoleWe're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on - and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently managing a large fleet of GPU-backed models.
We run nearly 30 models across heterogeneous hardware, and keeping that fleet fast, reliable, and cost-effective is a serious engineering challenge. You'll build the GPU management and scheduling platform that sits at the center of it - collecting utilization and load metrics, interpreting what they actually mean, and using them to make real-time decisions about admission control and scaling. The goal: route and schedule inference calls so we use our capacity efficiently without exceeding it, and scale model replicas up and down automatically as demand shifts.
This is a senior role for someone with a decade in the field who can move fluidly between systems engineering and software development, and who is excited to own a complex, evolving system end to end.
What You'll Do- Design and build our GPU management and scheduling platform - the system that decides when, where, and how inference calls run across a fleet of ~30 models on heterogeneous hardware
- Build the metrics pipeline that collects GPU load and utilization data, and the logic that turns those signals into decisions
- Implement admission control to protect capacity - deciding when to accept, queue, or shed inference requests so we operate within fleet limits
- Build autoscaling that adjusts the number of model replicas in response to real-time demand and utilization
- Develop cloud orchestration systems and operators in Python and Go to manage the model fleet
- Architect and operate scalable, fault-tolerant, secure production systems on AWS, GCP, or Azure
- Design and build infrastructure automation and deployment pipelines (Terraform, CI/CD) as first-class software
- Stand up and maintain monitoring, logging, and alerting that keep the platform reliable and performant
- Develop and enforce security and compliance policies appropriate to a healthcare AI platform
- Partner with engineers and research scientists to diagnose and resolve complex infrastructure, deployment, and operational issues
- Mentor engineers and raise the technical bar across the team
What You BringMust-Have
- 10+ years of professional experience across site reliability / DevOps engineering and software engineering
- Computer Science Degree Required from a top CS program.
- Strong software engineering fundamentals - you build orchestration and scheduling systems in Python and/or Go, not just configure off-the-shelf tools
- Experience designing systems that make decisions from operational metrics - collecting signals, interpreting them, and driving control loops such as autoscaling, load shedding, or admission control
- Deep experience with infrastructure automation and CI/CD (Terraform, GitLab CI/CD, or similar)
- Hands-on production experience with at least one major cloud platform (AWS, GCP, or Azure)
- Strong knowledge of containerization and orchestration (Docker, Kubernetes)
- Experience with monitoring and logging stacks (ELK, Grafana, Datadog, or similar)
- Familiarity with secrets management and security tooling (HashiCorp Vault, AWS KMS, Azure Key Vault)
- Excellent problem-solving skills and the ability to work both independently and collaboratively
- Strong communication and interpersonal skills
Nice-to-Have
- Experience managing GPU fleets or scheduling workloads across heterogeneous accelerators
- Familiarity with ML inference serving and model deployment (e.g. Triton, KServe, Ray Serve, or similar)
- Experience with Kubernetes autoscaling internals (HPA/VPA, custom metrics, custom controllers)
- Experience implementing HIPAA and SOC 2 compliance
- Experience operating in an HPC environment
- Bachelor's or Master's in Computer Science, Computer Engineering, or a related field
Join our team at Hippocratic AI and help shape the future of clinically safe, production-grade AI systems.