Manager - Production Operations & Site Reliability Engineering

Alcon

$140K — $181K *
Healthcare
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree or equivalent experience (high school + 13 years, Associate + 9 years, MS + 2 years, PhD + 0 years)
  • 5 years of relevant experience in Production Operations or Site Reliability Engineering
  • Preferred: Bachelor's degree in Computer Science, Engineering, or related field (Master's preferred)
  • Experience supporting regulated healthcare or medical device platforms, particularly SaaS applications
  • Certifications in AWS Solutions Architect or DevOps Professional, and Kubernetes (CKA/CKS)

Responsibilities

  • Lead steady-state operations for cloud-based healthcare platforms, ensuring high availability and performance.
  • Establish and improve operational standards, SLIs/SLOs, and service excellence practices.
  • Govern the Road to Production process, validating application readiness across environments.
  • Champion SRE best practices, leveraging automation to minimize operational toil.
  • Define enterprise observability strategies using tools like Datadog and CloudWatch.
  • Ensure compliance with HIPAA, GDPR, FDA, and cybersecurity standards.
  • Mentor engineering teams on operational excellence and influence architectural decisions.

Benefits

  • Collaborative work environment focused on innovation and personal growth.
  • Robust benefits package including health, life, and retirement options.
  • Opportunities for professional development and career advancement.
Full Job Description
Manager - Production Operations & Site Reliability Engineering

As a Principal Engineer you will provide technical leadership for the reliability, availability, security, and continuous improvement of Alcon's Digital Health Cloud platform supporting Software as a Medical Device (SaMD) and customer-facing digital health applications.

This role serves as the senior technical authority for production operations and Site Reliability Engineering (SRE), driving platform reliability, observability, automation, incident management, release governance, and cloud optimization. Partnering across Product Engineering, Architecture, Security, Infrastructure, and Operations, the Principal Engineer enables resilient, compliant, and scalable healthcare platforms with predictable, high-quality software delivery. In this role, a typical day will include:

Production Operations & Reliability
  • Lead steady-state operations for cloud-based healthcare platforms, ensuring high availability, reliability, and performance.
  • Establish and continuously improve operational standards, SLIs/SLOs, readiness reviews, and service excellence practices.
  • Drive platform resilience, capacity planning, disaster recovery, and governance through Alcon's Steady State Operations Framework (SSOF).

Release & Deployment Governance
  • Lead production readiness reviews, ensuring applications meet operational, security, monitoring, compliance, and supportability requirements.
  • Govern the Road to Production process across Validation, Staging, and Production environments, validating deployment readiness, infrastructure qualification, rollback plans, and acceptance criteria.
  • Partner with Engineering and DevOps teams to improve release quality, deployment reliability, and change success through standardized governance and automation.

Site Reliability Engineering (SRE)
  • Champion SRE best practices, leveraging automation and self-healing capabilities to reduce operational toil.
  • Improve service reliability and customer experience by reducing MTTD and MTTR and increasing deployment success rates.
  • Lead incident investigations, root cause analyses, and long-term corrective actions for critical production events.

Cloud Platform Operations
  • Provide technical leadership across AWS platforms including EKS, EC2, RDS, S3, ElastiCache, AWS MQ, Route53, and Kubernetes/Istio.
  • Optimize cloud infrastructure for scalability, resilience, security, performance, and cost efficiency.

Observability & Automation
  • Define enterprise observability strategies using Datadog, CloudWatch, distributed tracing, synthetic monitoring, centralized logging, and executive dashboards.
  • Lead automation initiatives across deployments, monitoring, health validation, incident response, and operational workflows.

Security & Compliance
  • Ensure compliance with HIPAA, GDPR, FDA, and enterprise cybersecurity standards.
  • Partner with Security teams to strengthen cloud security architecture, identity management, network segmentation, and operational controls.

Technical Leadership
  • Serve as the senior escalation point for major incidents, production events, and Hypercare operations.
  • Mentor engineering teams on operational excellence, production engineering, and SRE best practices.
  • Influence platform and architectural decisions that enhance operability, maintainability, resilience, and long-term service reliability.
  • Collaborate across Engineering, Architecture, Infrastructure, Security, and Global Operations to advance platform stability and operational maturity.


WHAT YOU'LL BRING TO ALCON:

Bachelor's Degree or Equivalent years of directly related experience (or high school +13 yrs; Assoc.+9 yrs; M.S.+2 yrs; PhD+0 yrs)

The ability to fluently read, write, understand and communicate in English

5 Years of Relevant Experience

PREFERRED QUALIFICATIONS:
  • Bachelor's Degree in degree in Computer Science, Engineering, or related field (Master's preferred)
  • Experience in Production Operations, Site Reliability Engineering, Platform Engineering, or Cloud Operations
  • Experience supporting regulated healthcare or medical device platforms
  • AWS Solutions Architect or DevOps Professional certification
  • Kubernetes (CKA/CKS) certification
  • Experience with healthcare interoperability standards (HL7, FHIR, DICOM)
  • Experience implementing Site Reliability Engineering practices within large-scale cloud environments
  • AWS (EKS, EC2, RDS, S3, Route53, IAM, Load Balancers, CloudWatch)
  • Kubernetes, Istio, Docker
  • Datadog, APM, logging, distributed tracing, synthetic monitoring
  • CI/CD, Infrastructure as Code, automation
  • HIPAA, GDPR, FDA compliance
  • Strong understanding of cloud networking, security, and high-availability architectures


HOW YOU CAN THRIVE AT ALCON:
  • Benefit from working in a highly collaborative environment.
  • Join Alcon's mission to provide top-tier, innovative products to enhance sight, enhance lives, and grow your career.
  • Alcon provides robust benefits package including health, life, retirement, PTO, and much more!


Alcon Careers

See your impact at alcon.com/careers

ATTENTION: Current Alcon Employee/Contingent Worker

If you are currently an active employee/contingent worker at Alcon, please click the appropriate link below to apply on the Internal Career site.

Find Jobs for Employees

Find Jobs for Contingent Worker

Compensation and Benefits

Alcon's Total Rewards programs are designed to align incentives with business objectives, support our values, and deliver long-term value. Our compensation approach includes a combination of fixed and variable pay, with short-term and long-term incentive opportunities for eligible roles. Our benefits offerings are designed to support associates and their families across key life events, including programs that promote health and well-being, provide financial security, and support retirement planning.

The salary range posted represents the anticipated hiring range for this role. Actual compensation may vary based on factors such as experience, skills, location, and internal equity, and may fall outside the posted range.

Pay Range
140,250.00 - 181,500.00

Pay Frequency
Annual

Similar Jobs

More Jobs at Alcon

More Healthcare Jobs

Find similar Manager - Production Operations & Site Reliability Engineering jobs: