Omnicell

Director, Site Reliability Operations

Omnicell$150K — $180K *
Information Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, IT, Engineering, or related field (Master's preferred)
  • 12+ years in cloud operations or related fields
  • 7+ years in progressive leadership managing operational teams
  • Experience leading 24×7 global production operations
  • Strong comprehension of cloud-native architectures and observability platforms
  • Proven background in operational transformation through automation
  • Exceptional executive communication skills.

Responsibilities

  • Build and lead Omnicell's Site Reliability Operations organization
  • Ensure continuous monitoring and management of cloud services
  • Own the Major Incident Management program
  • Establish operational governance for production events
  • Develop operational readiness standards for new services
  • Drive continuous improvement across production operations
  • Champion automation and AI-assisted operations.

Benefits

  • Remote or hybrid working options
  • Opportunity to redefine cloud operations
  • Access to cutting-edge technology and AI tools
  • Collaboration with cross-functional and executive teams
  • Focus on professional growth and operational excellence
  • Involvement in global cloud transformation initiatives.
Full Job Description
Job Description

Director, Site Reliability Operations (SRO)

Department: Global Cloud Operations
Reports To: Vice President, Global Cloud Operations
Location: Remote (U.S.) / Hybrid (Preferred)
Travel: Up to 20% (Domestic and International)

Site Reliability Operations (SRO) is the operational heartbeat of Omnicell's cloud platform. This organization is responsible for operating production services, maintaining situational awareness across the enterprise, coordinating incident response, restoring service, and continuously improving operational excellence.

As Director of Site Reliability Operations, you will build and lead a modern, AI-enabled operations organization that moves beyond the traditional Network Operations Center (NOC). Your team will leverage automation, observability, operational intelligence, and disciplined operational processes to proactively manage production environments and ensure exceptional service reliability for customers around the world.

This is an opportunity to redefine cloud operations by building a world-class Site Reliability Operations organization that partners closely with Site Reliability Engineering, Cloud Platform Engineering, Cloud Security, Product Engineering, and Technical Support.

About This Opportunity

This is not a traditional NOC leadership role.

Reporting directly to the Vice President of Global Cloud Operations, the Director of Site Reliability Operations will establish Omnicell's global operational strategy for cloud services, leading the teams responsible for 24×7 production operations, event management, service restoration, major incident management, operational governance, and operational readiness.

Working closely with the Director of Site Reliability Engineering, this organization will execute the operational strategy while SRE continuously engineers improvements to reliability, automation, and resilience. Together, these organizations create a modern cloud operating model where reliability is both engineered and operationally sustained.

The successful candidate will build a proactive, automation-driven operations organization that reduces operational risk, accelerates service restoration, and continuously improves the customer experience.

Purpose

The Director of Site Reliability Operations is responsible for leading Omnicell's global Site Reliability Operations organization.

This role owns the operational execution, governance, and continuous improvement of production cloud services, ensuring that customer-facing platforms remain available, resilient, and operationally efficient.

The organization serves as the central coordination point for operational awareness, incident response, service restoration, operational communications, and production readiness while driving automation that reduces manual effort and improves operational maturity.

Success in this role is measured through operational excellence, service availability, incident response effectiveness, customer experience, and continuous operational improvement.

Primary Impact

As Director of Site Reliability Operations, you will define how Omnicell operates its production cloud environment.

Your leadership will establish a proactive operational culture that leverages automation, observability, AI-assisted operations, and disciplined operational processes to detect issues early, minimize customer impact, accelerate service restoration, and continuously improve operational performance.

Working alongside Site Reliability Engineering, Cloud Platform Engineering, and Cloud Security, your organization will ensure that production systems remain stable, resilient, and ready to support Omnicell's continued growth.

What You'll Do

Build and Lead a World-Class Site Reliability Operations Organization

Build and scale Omnicell's global Site Reliability Operations organization.

You will:
  • Recruit, mentor, and develop Operations Managers and Site Reliability Operations Engineers.
  • Build a high-performing global operations organization.
  • Establish operational standards, career development frameworks, and leadership expectations.
  • Foster a culture of ownership, accountability, continuous improvement, and operational excellence.
  • Develop a follow-the-sun operational model supporting global cloud services.


24×7 Production Operations

Lead Omnicell's enterprise production operations organization responsible for continuous monitoring and operational management of cloud services.

Responsibilities include:
  • Production Operations
  • Operational Monitoring
  • Service Health Management
  • Event Management
  • Alert Triage
  • Operational Escalation
  • Shift Operations
  • Operational Command Center
  • Customer Impact Assessment
  • Operational Coordination

Ensure operational teams maintain continuous awareness of platform health while proactively identifying and mitigating production risks.

Major Incident Management

Own Omnicell's enterprise Major Incident Management (MIM) program.

Responsibilities include:
  • Executive Incident Command
  • Major Incident Coordination
  • Cross-Functional War Rooms
  • Service Restoration
  • Stakeholder Communications
  • Executive Communications
  • Customer Communications (in partnership with Support)
  • Incident Documentation
  • Incident Timeline Management
  • Escalation Governance

Lead high-severity incidents with urgency, structure, and transparency while minimizing business and customer impact.

Event & Operational Management

Establish enterprise operational governance for production events.

Own processes supporting:
  • Event Correlation
  • Alert Quality
  • Noise Reduction
  • Event Prioritization
  • Operational Dashboards
  • Monitoring Effectiveness
  • Operational KPIs
  • Operational Health Reviews

Partner with Site Reliability Engineering to continuously improve monitoring quality and operational intelligence.

Operational Readiness

Develop operational readiness standards for new services entering production.

Responsibilities include:
  • Operational Readiness Reviews
  • Runbook Validation
  • Playbook Development
  • Operational Acceptance
  • Monitoring Validation
  • Escalation Readiness
  • Support Readiness
  • Disaster Recovery Readiness
  • Service Transition

Partner with Product Engineering and Site Reliability Engineering to ensure services are operationally prepared before production deployment.

Operational Excellence & Continuous Improvement

Drive continuous improvement across production operations.

Lead initiatives focused on:
  • Reducing MTTD (Mean Time to Detect)
  • Reducing MTTR (Mean Time to Restore)
  • Improving Incident Quality
  • Eliminating Recurring Operational Issues
  • Standardizing Operational Processes
  • Increasing Automation Adoption
  • Improving Service Availability
  • Enhancing Customer Experience

Establish a culture that uses operational metrics and post-incident learning to drive measurable improvements.

Operational Automation & AIOps

Champion automation and AI-assisted operations across the production environment.

Lead initiatives supporting:
  • Automated Incident Response
  • Self-Healing Workflows
  • Event Correlation
  • Intelligent Alerting
  • Automated Runbooks
  • Operational ChatOps
  • AI-Assisted Root Cause Analysis
  • Predictive Operations
  • Operational Knowledge Management

Partner with Site Reliability Engineering and Cloud Platform Engineering to engineer automation that reduces operational toil and improves operational consistency.

Operational Governance & Service Management

Provide governance across enterprise operational processes.

Own:
  • Operational Policies
  • Incident Governance
  • Problem Management
  • Change Coordination
  • Service Health Reporting
  • Operational Risk Reviews
  • SLA Compliance
  • Operational Metrics
  • Executive Operational Reporting

Ensure consistent operational practices across all production environments.

Managed Service Provider (MSP) Governance

Lead operational governance for Omnicell's strategic managed service providers.

Responsibilities include:
  • Operational Performance Management
  • SLA Governance
  • Vendor Escalation Management
  • Operational Reviews
  • Performance Scorecards
  • Continuous Improvement Programs
  • Contractual Operational Alignment

Ensure third-party operational partners consistently meet Omnicell's operational standards and customer expectations.

Cross-Functional Partnership

Develop trusted partnerships across:
  • Site Reliability Engineering
  • Cloud Platform Engineering
  • Cloud Security
  • Product Engineering
  • Technical Support
  • Enterprise Architecture
  • Customer Success
  • Engineering Leadership

Serve as the operational bridge between engineering organizations and customer-facing support teams.

Organizational Leadership

Provide strategic leadership by:
  • Building a globally distributed Site Reliability Operations organization.
  • Developing future operational leaders.
  • Driving employee engagement and professional growth.
  • Creating a culture centered on operational discipline and customer focus.
  • Encouraging innovation and operational excellence.


Executive Leadership

Partner with executive leadership to:
  • Present operational health and service performance.
  • Recommend investments that improve operational maturity.
  • Communicate production risks and mitigation strategies.
  • Lead executive incident communications during major outages.
  • Support enterprise cloud transformation initiatives.
  • Influence enterprise operational strategy.


What Success Looks Like

First 90 Days
  • Assess current production operations and operational maturity.
  • Build relationships with Engineering, Product, Security, and Support leaders.
  • Evaluate monitoring, incident management, and operational processes.
  • Identify opportunities to improve operational effectiveness.
  • Develop a multi-year Site Reliability Operations strategy.


First Six Months
  • Standardize enterprise incident management processes.
  • Improve operational dashboards and executive reporting.
  • Implement operational governance standards.
  • Launch automation initiatives to reduce manual operational effort.
  • Recruit key operational leadership positions.
  • Establish operational KPIs and scorecards.


First Year
  • Build a high-performing global Site Reliability Operations organization.
  • Improve MTTD, MTTR, and overall service availability.
  • Reduce alert fatigue through improved monitoring quality.
  • Expand operational automation and AI-assisted operations.
  • Establish a mature, scalable operational model supporting Omnicell's global cloud platform.
  • Become a trusted operational partner across Engineering, Product, Support, and Executive Leadership.


Who You Are

You are a servant leader who thrives in dynamic production environments and believes operational excellence is built through discipline, teamwork, automation, and continuous learning.

You excel at coordinating complex operational events, building high-performing operations teams, and transforming reactive operational processes into proactive, data-driven capabilities.

You understand that exceptional cloud operations require strong partnerships with engineering organizations, customer-facing teams, and executive leadership. Most importantly, you are passionate about creating resilient operational systems that enable engineers to innovate while ensuring customers experience reliable, always-available services.

Minimum Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field (Master's degree preferred).
  • 12+ years of experience in cloud operations, production operations, NOC, SRE, IT Operations, or related disciplines.
  • 7+ years of progressive leadership experience managing managers and operational teams.
  • Experience leading 24×7 global production operations organizations.
  • Strong understanding of cloud-native architectures, observability platforms, incident management, and operational governance.
  • Experience driving operational transformation through automation and process improvement.
  • Exceptional executive communication and crisis leadership skills.


Preferred Qualifications
  • Experience building modern Site Reliability Operations or cloud operations organizations.
  • Experience with AIOps, observability platforms, and operational analytics.
  • Expertise with ITIL, Incident Management, Problem Management, and Change Management frameworks.
  • Experience supporting regulated healthcare, SaaS, or enterprise cloud platforms.
  • Demonstrated success leading globally distributed operations organizations.
  • Experience managing strategic managed service providers.


Leadership Expectations

As Director of Site Reliability Operations, you will:
  • Establish the operational vision for Omnicell's production

About Omnicell

Omnicell, Inc. is an American multinational healthcare technology company headquartered in Mountain View, California. It manufactures automated systems for medication management in hospitals and other healthcare settings, and medication adherence packaging and patient engagement software used by retail pharmacies. Its products are sold under the brand names Omnicell and EnlivenHealth.
Learn more about Omnicell
Size
3,800 employees
Market Cap
$2 billion
Industry
Net Income
$32.1 million
Founded
1992
5 Year Trend
+10.2%
Revenue
$892.2 million
NASDAQ

Similar Jobs

More Jobs at Omnicell

More Information Technology Jobs

Find similar Director, Site Reliability Operations jobs: