Job DescriptionSite Reliability Engineer (SRE) Location: Hybrid (U.S.)
Department: Global Cloud Operations
Reports to: VP, Global Cloud Operations
What You'll Do Primary Impact: As a Site Reliability Engineer, you will own the reliability, scalability, and operational health of a defined set of cloud services that support mission-critical pharmacy automation systems used by healthcare providers worldwide.
Service Reliability & Automation
- Own reliability outcomes for assigned services, ensuring strong instrumentation, actionable alerts, meaningful dashboards, and up-to-date runbooks.
- Define and implement SLIs and SLOs in partnership with product and engineering teams, and surface reliability performance in regular Cloud Operations reviews.
- Identify operational toil and design automation to eliminate repetitive manual work.
- Drive continuous improvement initiatives that increase observability, automation coverage, and system resilience.
Incident Response & Operational Excellence
- Participate in the SRE on-call rotation, progressing from secondary to primary ownership as readiness increases.
- Command Sev-2 and Sev-3 incidents independently over time, with pairing and coaching from a Senior SRE; act as technical lead during Sev-1 incidents.
- Lead blameless post-incident reviews and own follow-up actions through completion.
- Partner closely with managed services providers (IBM, HCL) to ensure clean escalation paths from L1/L2 monitoring into SRE ownership.
Platform, CI/CD & Observability
- Design, build, and operate CI/CD pipelines supporting cloud-native application delivery using tools such as GitHub Actions, CodeFresh, TeamCity, and Octopus Deploy.
- Automate infrastructure and platform services using Infrastructure as Code (Terraform preferred).
- Contribute to the evolution of Omnicell's observability platform, including intelligent alerting, ML-based anomaly detection, and automated diagnostics.
- Participate in architecture and launch readiness reviews, bringing a reliability lens to system design.
- Help establish reference implementations and "golden paths" that enable product teams to launch services with reliability built in from day one.
Who You Are - Bachelor's degree in Computer Science, Engineering, or a related technical field.
- 5+ years of experience in software or platform engineering, including 3+ years in an SRE, DevOps, or reliability-focused role.
- Strong hands-on experience with at least one major public cloud platform (AWS, Azure, or GCP).
- Proficiency in Python or another object-oriented programming language for automation and tooling.
- Production experience with Kubernetes, Docker, and Helm.
- Experience implementing Infrastructure as Code using Terraform or similar frameworks.
- Working knowledge of modern observability tools across metrics, logs, and tracing.
- Real-world incident response experience, including on-call participation and post-incident write-ups.
- Solid Linux system administration skills.
- Collaborative, coachable mindset with a desire to grow under senior mentorship.
Preferred Qualifications - Experience working in regulated environments such as healthcare, financial services, or government (HIPAA, SOC 2, or similar).
- Familiarity with managed service provider models for L1/L2 operations.
- Exposure to AIOps, ML-based anomaly detection, or LLM-assisted incident triage.
- Understanding of GitOps principles and tools such as ArgoCD or Flux.
- Experience operating secure, compliant Kubernetes platforms.
- Familiarity with chaos engineering, messaging systems (Kafka, RabbitMQ), or stateful services in Kubernetes.
How You'll Elevate at Omnicell At Omnicell, success isn't just about what you deliver-it's about
how you deliver it. Our Elevate Behaviors guide how we work together and create impact:
- Collaborate: Partner closely with product engineering, security, and operations teams to build shared ownership of reliability.
- Inspire: Influence reliability best practices across teams by modeling calm, structured incident leadership.
- Develop: Continuously build your technical depth while learning directly from a senior SRE mentor.
- Execute: Take ownership of services, incidents, and follow-through-turning lessons learned into measurable improvements.
- Impact: Help shape foundational SRE practices and introduce modern reliability and AIOps capabilities that scale with the business.
Growth & Career Path This role is intentionally designed as a
growth role. With strong performance and increasing ownership, the natural progression is into a
Senior Site Reliability Engineer position as the practice scales. Omnicell also supports lateral growth into platform engineering, security engineering, or product engineering for SREs who discover adjacent passions.
Work Conditions - Remote or hybrid work environment supported.
- Up to 10% travel as needed.
- Participation in an SRE on-call rotation is required.