Site Reliability Engineering Team Lead (Principal SRE)

Cerence Inc.$150K — $180K *
US-Anywhere
+ 2 other locationsRemote
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in site reliability, DevOps, or cloud roles, with leadership experience
  • Ability to set technical direction and uphold standards across teams
  • Hands-on experience with Kubernetes, Docker, and Istio
  • Familiarity with Azure, AWS, and Google Cloud platforms
  • Knowledge of observability tools like Zabbix, Prometheus, and Grafana
  • Proficient in CI/CD practices using Terraform or Flux
  • Strong skills in scripting languages like Python or Go
  • Excellent communication skills in English.

Responsibilities

  • Lead the Site Reliability Engineering team with technical direction and mentoring.
  • Own the reliability roadmap for upcoming quarters.
  • Define and manage SLI/SLO standards for customer program availability.
  • Serve as Tier 2 escalation for major incidents, managing high-stakes situations.
  • Promote and implement a blameless postmortem culture for incident management.
  • Enhance production readiness by collaborating with development teams.
  • Drive metrics and automation strategies for observability and CI/CD.

Benefits

  • Annual bonus opportunity
  • Comprehensive insurance coverage including medical, dental, vision, and life insurance
  • Paid time off and holidays
  • Company contributions to a Registered Retirement Savings Plan (RRSP)
  • Equity awards available for certain positions
  • Remote or hybrid work options based on position.
Full Job Description
Site Reliability Engineering Team Lead (Principal SRE)
Job Description Summary

The Site Reliability Engineering Team Lead (Principal SRE) leads Cerence's Site Reliability Engineering team, owning the reliability, availability, and operational health of our cloud-native automotive AI platform. This role combines technical leadership of the team and the function with deep technical credibility: you will help select and mentor the team, define the reliability roadmap, govern SLI/SLO/SLA targets, and champion a blameless culture. You will serve as the Tier 2 technical escalation point for major incidents, partner with development leadership to embed reliability into the SDLC, and set strategic direction for observability and automation. This is a Principal-level role on our technical track, with no direct reports: you lead the team and own the function through technical authority. The ideal candidate brings a strong hands-on background in SRE, cloud platforms, and container orchestration, and a track record of leading site reliability teams.

What we offer
  • We offer a generous compensation and benefits package (in addition to the base salary), including:
  • Annual bonus opportunity
  • Insurance coverage (medical, dental, vision, life, and disability)
  • Paid time off
  • Paid holidays
  • Company contribution to the RRSP (Registered Retirement Savings Plan)
  • Equity awards for certain positions and levels
  • Remote and/or hybrid work available depending on the position All compensation and benefits are subject to the terms and conditions of the underlying plans or programs, as applicable, and may be amended, terminated, or replaced from time to time.

The Opportunity

We are looking for an experienced Principal SRE to lead our Site Reliability Engineering team and own the reliability function. In this role, you will own the reliability of Cerence's cloud-native voice, gesture, and gaze solutions - systems that power immersive automotive AI experiences at global scale. The function is yours end to end: the reliability roadmap, the SLI/SLO standards, the on-call model, the Tier 2 escalation path, and approval authority over high-risk production change. You will set the team's technical direction, build a sustainable team culture, and be the bridge between reliability practice and product delivery.

This is a hands-on leadership role. You will bring both leadership experience and deep technical credibility, enabling you to earn the trust of your team, hold informed architectural conversations, and make sound trade-offs under pressure.
Principal Responsibilities

Technical Leadership
  • Help select, mentor, and technically develop the team, across multiple locations
  • Set technical direction and priorities for the team, and contribute performance and growth input to their managers
  • Design and maintain a sustainable on-call rotation; monitor page load and team health as first-class concerns

Reliability Strategy
  • Own and drive the team's reliability roadmap across a 2-3 quarter horizon
  • Define and govern SLI/SLO/SLA frameworks to hold the contracted availability targets for our customer programs, which run as high as 99.95%

Incident & Production Excellence
  • Serve as the Tier 2 technical escalation point for major incidents, partnering with the Global Operations Center, which owns incident management and response
  • Champion blameless postmortem culture - model it, reinforce it, and ensure it produces actionable outcomes
  • Lead and continuously improve Production Readiness / NFR reviews with development teams
  • Contribute to root cause analysis and own the systemic improvements that come out of it
  • Act as a named approver for high-risk and out-of-window production change

Observability & Automation
  • Set the strategic direction for metrics, dashboards, and alerting across SLI/SLO, escalation, and automation layers
  • Drive development and adoption of CI/CD automation pipelines for service deployments, rollbacks, and operational tasks
  • Partner with DevOps and platform teams to evolve shared infrastructure

Cross-Functional Engagement
  • Partner with development managers and architects to embed reliability into the SDLC by default
  • Participate in service reliability consulting and architectural reviews
  • Communicate reliability posture and risk clearly to technical and non-technical stakeholders
Required Qualifications
  • 8+ years of hands-on experience in site reliability, DevOps, or cloud platform roles, including time leading a team or owning a function
  • A track record of setting technical direction and holding standards across a team - with or without formal authority
  • Hands-on experience with container orchestration frameworks (Kubernetes, Docker, Istio)
  • Experience with public cloud platforms (Azure primarily; AWS and Google Cloud)
  • Familiarity with observability tooling - metrics pipelines, dashboarding, and alerting (e.g., Zabbix, Prometheus, Grafana)
  • Experience with CI/CD pipelines and infrastructure-as-code practices (e.g., Terraform, Flux)
  • Proficiency in at least one scripting or programming language (Python, Go, Shell, etc.)
  • Strong UNIX/Linux background, including system configuration, performance debugging, and network fundamentals (Layer 4/5, DNS, HTTP/S, TLS)
  • Excellent written and verbal communication skills in English
Preferred Qualifications
  • Previous site reliability leadership experience
  • Experience leading distributed or multi-site technical teams
  • Background in high-availability service design (redundancy, failover, blast radius)
  • Experience with log aggregation and analytics platforms (Loki, Thanos)
  • Familiarity with ITSM and project tooling (Jira, Confluence)
  • Experience in automotive, embedded, or latency-sensitive production environments

About Cerence Inc.

Cerence Inc. is a software company that specializes in voice recognition and natural language understanding technology. The company was spun off from Nuance Communications in 2019 and is headquartered in Newton, Massachusetts. Cerence's software is used in a variety of applications, including automotive infotainment systems, smart speakers, and virtual assistants. The company's clients include many of the world's leading automakers, as well as companies in the consumer electronics and mobile device industries. Cerence has received several awards for its technology, including the 2020 CES Innovation Award for its Cerence Drive platform.
Learn more about Cerence Inc.
Size
1,200 employees
Market Cap
$726.1 million
Industry
Net Income
$12.7 million
5 Year Trend
+6%
Revenue
$347.1 million

Similar Jobs

More Jobs at Cerence Inc.

More Information Technology Jobs

Find similar Site Reliability Engineering Team Lead (Principal SRE) jobs: