Eli Lilly

Senior Principal SRE Engineering

Eli Lilly$129K — $231K *
Pharmaceuticals & Biotech
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science or related technical discipline.
  • 7+ years of progressive engineering experience, including at least 6 as a Site Reliability Engineer.
  • Hands-on experience with enterprise-scale SRE platforms and observability solutions.
  • Proven ability in designing and implementing self-healing patterns in production environments.
  • Experience creating and managing SLO/SLI frameworks for multi-application estates.
  • Strong technical leadership and mentoring experience with senior engineers.

Responsibilities

  • Define service-level objectives and govern error-budget burn across applications.
  • Establish observability and instrumentation standards for all applications.
  • Drive blameless postmortems and ensure root causes lead to engineering work.
  • Ensure compliance with Lilly standards and promote operational security.
  • Set the engineering bar and mentor senior reliability engineers.

Benefits

  • Comprehensive medical, dental, and vision benefits.
  • 401(k) and pension plan participation.
  • Flexible spending accounts for health and childcare.
  • Life insurance and death benefits.
  • Employee assistance and wellness programs.
Full Job Description
About the Team

Technology at Lilly builds and maintains capabilities using pioneering technologies like the most prominent tech companies. What differentiates Lilly IT is that we redefine what's possible through tech to advance our purpose, creating medicines that make life better for people around the world, including data-driven drug discovery, connected clinical trials, resilient enterprise platforms, and intelligent digital operations. We hire the best technology professionals from a variety of backgrounds, so they can bring an assortment of knowledge, skills, and diverse thinking to deliver creative solutions in every area of our business.

The Digital Core team leads Lilly's transformation into the Digital and AI era. They inspire digitally empowered teams to new ways of working and accelerate innovation and agility. This team powers and advances the entire company by building and maintaining world-class technology capabilities and platforms.

The Reliability Engineering team is the engineering-first function that owns the stability, observability, and operational quality of a multi-application production estate. It operates in close partnership with the engineering team that builds the agentic automation platform, and is in active transition from human-executed operations to engineering-led, agent-assisted reliability.

Role Summary:
As the Senior Principal SRE Engineering Lead, you are the senior-most engineering authority for the reliability of the supported production estate. You own the bar for what reliable means in this organization: service-level objectives, error-budget governance, observability standards, and the engineering practices that protect both production and the team's engineering time.

This role is the judgment layer between an agent's recommendation and a production change. You combine engineering rigor with operational pragmatism, and you decide which patterns surfaced from production warrant durable engineering investment. You hold the SRE engineering bar within the function - the cross-pillar reference architecture is owned by the Senior Architect, but the practice, standards, and engineering judgment of reliability at this site are yours.

You are a senior individual contributor. You do not manage people. You partner with the Reliability Leader, the Senior Principal Tech Shift Lead in the same pillar, the Senior Architect, and the senior engineering principals in the agentic automation team. Success is measured by organizational impact, sustained reliability outcomes, and the ability to scale reliability through systems and people, not heroics.

What you'll be doing:

1) Reliability strategy and SLO governance
  • Define service-level objectives and indicators across the supported production estate, tiered by application risk and business impact.
  • Govern error-budget burn: when to slow change, when to invest in durable fixes, when to accept the budget.
  • Drive the adoption of reliability reporting and the disciplined operating cadence that makes SLOs real, not decorative.
  • Hold the SLO and error-budget conversation with product and application owners - including the conversation about what their service needs to change to meet the bar.

2) Engineering standards, observability, and self-healing
  • Establish observability and instrumentation standards as the contract every supported application must meet - and hold the line on them.
  • Set the bar for infrastructure-as-code, continuous-delivery hardening, and deployment safety across the supported estate.
  • Define the engineering work that flows from incidents and root-cause analyses into durable production change, and the architectural patterns (self-healing runbooks, graceful degradation, circuit breakers) that reduce repeat failure.

3) Incident learning, durable fixes, and partnership with agentic automation
  • Drive blameless postmortem culture; ensure root-cause analyses produce engineering work, not just narrative.
  • Govern which patterns from production warrant durable engineering investment, against the error-budget regime.
  • Partner with the production operations team in the same pillar on which patterns from the field warrant engineering attention; partner with the agentic automation engineering team on which fixes become safe agent-assisted remediations, and on the confidence thresholds, guardrails, and human-in-the-loop boundaries that make those remediations safe in production.
  • Lead high-severity incident response as incident commander when escalation reaches this seat, and coach the team to handle the rest.

4) Compliance, security, and regulated-environment readiness
  • Ensure reliability practices comply with Lilly standards and applicable regulatory requirements, and engineer them so that audit evidence falls out of normal operation.
  • Promote secure operational practices, auditability, and validated-environment-friendly engineering as part of the standards the team holds.
  • Act as a trusted technical leader in regulated and validated environments.

5) Technical leadership and talent development
  • Set the engineering bar through standards, expectations, and role modeling.
  • Mentor senior reliability engineers in the pillar; build the bench for sustained team growth and develop the next layer of principal engineers.
  • Influence engineering, product, and platform leaders through credibility and outcomes rather than authority.
  • Contribute to the evolution of enterprise-wide reliability practices in partnership with the Senior Architect and peer technical leaders.

How you will succeed

At the senior-most engineering individual-contributor level for reliability, success is defined by breadth of impact and sustained outcomes:
  • Be recognized as the senior reliability authority for your area.
  • Demonstrate measurable, sustained improvements such as: reduced major incidents, fewer recurring failures, improved time-to-recovery, and a credible error-budget regime.
  • Influence decisions across multiple teams and leaders through expertise and trust.
  • Scale reliability through systems, standards, and people, not heroics.


Your Basic Qualification:
  • Bachelor's degree in Computer Science, Information Technology, or a related technical engineering discipline, including Software Engineering, Computer Engineering, Information Systems, Cybersecurity, Information Science, Network Engineering, Systems Engineering, Computer Information Systems (CIS), Management Information Systems (MIS), Cloud Computing, Data Science
  • 7+ years of progressive engineering experience, with at least 6 years as a Site Reliability Engineer, Production Engineer, or equivalent, including a tour as the senior-most reliability engineer for a multi-application production estate - not a single product.
  • Hands-on experience authoring and validating runbooks: safe execution order, rollback steps, and exception handling for real remediation procedures. Demonstrated hands-on experience designing, implementing, and operating enterprise-scale SRE platforms, including observability solutions (Splunk, Datadog, New Relic, or Grafana/Prometheus), infrastructure-as-code with Terraform, CI/CD pipeline hardening, Kubernetes-based container platforms, and production workloads hosted on AWS, Azure, or GCP.
  • Experience designing self-healing patterns (circuit breakers, graceful degradation, automated remediation) and validating them before they're trusted in production. Production reliability experience in a regulated or audited environment (GxP, SOX, HIPAA, PCI, or equivalent), including familiarity with change-control discipline, audit evidence, and validated-system constraints.
  • Hands-on ownership experience of an SLO/SLI framework and error-budget policy across a multi-application estate, including defining SLIs, negotiating SLOs with product owners, and operationalizing burn-rate alerting and error-budget governance.
  • Experience leading high-severity incident response as incident commander, running blameless postmortems, and converting findings into durable engineering work - measured by reduced recurrence rather than narrative quality.
  • Demonstrated technical leadership experience at scale: mentoring senior ICs, influencing engineering and product leaders without org-chart authority, and writing the standards and engineering documents that set the bar for a discipline.


What You Should Bring:
  • Hands-on experience designing self-healing automation and running chaos engineering or resilience-testing programs (AWS Fault Injection Service, Gremlin, LitmusChaos, or equivalent) tied to measurable reliability gains.
  • Deep AWS fluency across reliability-relevant services (EKS, ECS, Lambda, CloudWatch, X-Ray, Systems Manager, Route 53), and familiarity with AWS Well-Architected Reliability Pillar.
  • Experience with AIOps or agent-assisted operations, including designing the guardrails, confidence thresholds, and human-in-the-loop boundaries that make automated remediation safe in production.
  • Track record of saying no to a deploy because the error budget was burned, and making the call stick.
  • Prior experience building a reliability practice from a small founding team.
  • Experience operating in pharma, healthcare, financial services, or other regulated industries.


Leadership expectations
  • Acts as the judgment layer between aspiration and engineering reality.
  • Combines engineering rigor with operational pragmatism, and knows when each is the right answer.
  • Leads through what they build and how they write, not through org-chart authority.
  • Comfortable telling a product owner that their application does not meet the reliability bar yet, and showing them how to get there.
  • Treats mentorship of senior reliability engineers as a first-class outcome of the role.

Additional information

Availability to work flexible work hours is/may be required. This team supports continuous operations and may require non-standard work hours, including some work on weekends and holidays.

Actual compensation will depend on a candidate's education, experience, skills, and geographic location. The anticipated wage for this position is
$129,000 - $231,000

Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly's compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.

#WeAreLilly

About Eli Lilly

ICOS Corporation is a biotechnology company that engages in the discovery, development, and commercialization of therapeutic products. It is engaged in the commercialization of treatments for unmet medical conditions, such as benign prostatic hyperplasia, hypertension, pulmonary arterial hypertension, cancer, and inflammatory diseases. It is the developer of a treatment known as Cialis (tadalafil), a product for the treatment of erectile dysfunction through its joint venture with Eli Lilly and Company in North America and Europe. It is also engaged in contract manufacturing services for third parties. It is in a strategic alliance with Solvay Pharmaceuticals, Inc. ICOS Corporation was established in 1989, based in Bothell, Washington. It is currently operated by Eli Lilly and Company.

Eli Lilly Careers

Joining Eli Lilly offers an unparalleled opportunity to become part of a leading global team dedicated to creating a healthier future. As a company revered for its commitment to innovation and leadership in the pharmaceutical industry, Eli Lilly is where your professional journey can flourish. Work You’ll Do At Eli Lilly, we are passionate about transforming patient care and advancing medical innovation. Our team at Eli Lilly is at the forefront of developing groundbreaking solutions in healthcare. By joining us, you will collaborate with some of the brightest minds in the industry, using cutting-edge technology to make real-world impacts. Lead with Innovation and Leadership Eli Lilly stands out in the marketplace by integrating deep industry expertise with robust research and development efforts. We are looking for professionals who are eager to drive change and lead the way in developing therapeutic breakthroughs. Explore Job Opportunities and Growth Eli Lilly offers a variety of career paths, including full-time positions and internships, across multiple functions such as research, marketing, IT, and sales. Whether you are a seasoned professional or a recent graduate, Eli Lilly provides an environment that promotes career growth and learning opportunities. Our commitment to diversity and leadership training ensures that every employee can achieve their potential. Be Part of Our Team Our team at Eli Lilly is committed to excellence and driven by a mission to improve lives. Employees enjoy a supportive culture that values collaboration, creativity, and diversity. We believe that a diverse workforce fosters innovation and helps us better connect with the communities we serve. Benefits and Culture Eli Lilly is dedicated to supporting our employees, offering competitive benefits, wellness programs, and comprehensive health care. Our culture is built on a foundation of respect, integrity, and quality, making Eli Lilly not just a great place to work, but a community to grow with. Networking and Professional Development Eli Lilly encourages continuous professional development and networking. With access to various training programs and mentorship opportunities, employees can enhance their skills and advance their careers. Our leadership is committed to nurturing talent through effective training and development strategies. Join Our Team Discover the exciting job opportunities at Eli Lilly by exploring open positions that match your skills and interests. We are continuously hiring and looking for individuals who are passionate, innovative, and ready to contribute to our mission of making life better for people around the globe. Stay Connected Keep up to date with the latest at Eli Lilly by following our careers blog. Gain insights from industry leaders and get tips on everything from crafting the perfect resume to preparing for your interview. Eli Lilly is not just a company—it's a place where you can make a difference. Explore the positions available and find out how your talents can help change the world. SEARCH ELI LILLY JOBS Stay ahead in your career with Eli Lilly, where innovation, leadership, and a commitment to diversity and growth lead the way to future advancements.
Learn more about Eli Lilly
Size
35,000 employees
Market Cap
$344.2 billion
Industry
Net Income
$6.1 billion
Founded
1876
5 Year Trend
+5.9%
Revenue
$24.5 billion
NASDAQ

Similar Jobs

More Jobs at Eli Lilly

More Pharmaceuticals & Biotech Jobs

Find similar Senior Principal SRE Engineering jobs: