Senior Manager, Incident Management

Zillow, Inc.

$125K — $211K *
US-AnywhereRemote in United States
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in incident management, SRE, or technical operations with 2+ years in people management
  • Experience in building or scaling incident and problem management programs
  • Strong knowledge of incident management practices and problem management methodologies
  • Proficiency in AI tools to enhance operational processes
  • Exceptional communication skills for technical and executive audiences
  • Ability to maintain composure under pressure during critical incidents
  • Skilled in coaching and developing operational talent

Responsibilities

  • Lead the incident management program across the organization from process to governance
  • Act as the executive escalation point for major incidents
  • Establish standards for incident severity classification and communication protocols
  • Collaborate with leadership to align incident management with business goals
  • Drive communications regarding incidents at the executive level
  • Develop and manage a formal problem management practice
  • Ensure root cause analyses for incidents include candid evaluations and actionable items
  • Identify and remediate recurring issues and trends from past incidents
  • Promote AI-driven workflows to enhance incident response speed and quality
  • Mentor and develop a team of incident managers while ensuring a healthy team structure
  • Track metrics to measure program effectiveness and drive continuous improvement

Benefits

  • Eligible for equity awards based on experience and performance
  • Opportunity for career growth and development within the organization
  • Collaborative work environment with cross-functional partnerships
  • Participation in shaping the organization’s approach to AI and automation in operations
  • Engagement in a culture of continuous improvement and blameless learning
Full Job Description
About the role
Responsibilities
Incident Management Leadership
  • Own the end-to-end incident management program, including process design, tooling, and governance across the organization
  • Serve as executive escalation point and senior decision-maker for critical, high-severity, or cross-functional incidents
  • Set and enforce standards for incident severity classification, escalation paths, and communication protocols
  • Partner with Engineering, Product, and business leadership to align incident response with business priorities and risk tolerance
  • Drive executive-level incident communications, ensuring leadership has clear, timely, and accurate visibility into impact and status
Problem Management
  • Build and own a formal problem management practice that connects incident trends to systemic root causes
  • Ensure every significant incident produces a rigorous, blameless root cause analysis (RCA) with clearly owned, tracked corrective actions
  • Establish mechanisms to identify recurring issues, chronic risks, and process gaps across incident history
  • Hold cross-functional partners accountable for closing problem records and remediation items on committed timelines
  • Report on problem management outcomes and reliability trends to leadership, tying them to measurable risk reduction
Leveraging AI Workflows for Speed & Quality
  • Champion the adoption of AI-powered tooling and workflows across incident detection, triage, summarization, and RCA drafting
  • Design and continuously improve AI-assisted workflows that turn raw incident and problem data into clear, actionable insights for internal partners
  • Ensure AI-generated call-outs, summaries, and reports meet a high bar for accuracy, relevance, and actionability before reaching stakeholders
  • Identify opportunities to automate repetitive operational tasks (documentation, status updates, trend analysis) to free the team to focus on higher-value judgment work
  • Partner with Engineering and Data teams to pilot, evaluate, and scale new AI capabilities that improve mean-time-to-resolution and mean-time-to-detection
People & Team Leadership
  • Hire, coach, and develop a team of incident managers, building depth and bench strength across severity levels
  • Set clear performance expectations and career growth paths for the team
  • Establish on-call structures, workload balance, and rotations that sustain team health and reliability coverage
  • Foster a culture of ownership, continuous improvement, and blameless learning within the team
Operational Excellence & Continuous Improvement
  • Define and track key metrics (MTTR, MTTD, recurrence rate, action-item closure rate) to measure program health and impact
  • Continuously refine runbooks, tooling, and workflows based on retrospectives and data trends
  • Facilitate post-incident reviews for major incidents, ensuring lessons learned translate into concrete process or system changes
  • Benchmark practices against industry standards and bring in outside best practices where relevant
Scope & Impact
  • Owns the incident and problem management strategy and roadmap for the organization
  • Accountable for outcomes across the full incident lifecycle, from detection through remediation and prevention
  • Directly manages a team of incident managers and indirectly influences engineering and support teams during incident response
  • Shapes how AI and automation are applied across operational workflows, with measurable impact on speed and quality of output
  • Decisions and process changes influence reliability posture and stakeholder trust across the broader organization


This role has been categorized as an Office position. "Office" employees regularly work at an existing ZG corporate office for approximately 80 to 100 percent of their time each month. Employees must live within reasonable commuting distance of their designated ZG office. ZG has not defined a reasonable distance, and expects employees will use judgment in determining this for themselves and understand the implications re: time commitment and cost of daily commute.

In California, Connecticut, Maryland, Massachusetts, New Jersey, New York, Washington state, and Washington DC the standard base pay range for this role is $132,400.00 - $211,600.00 annually. This base pay range is specific to these locations and may not be applicable to other locations.In Colorado, Hawaii, Illinois, Maine, Minnesota, Nevada, Ohio, Rhode Island, Vermont, and Virginia the standard base pay range for this role is $125,800.00 - $201,000.00 annually. The base pay range is specific to these locations and may not be applicable to other locations.

In addition to a competitive base salary this position is also eligible for equity awards based on factors such as experience, performance and location. Actual amounts will vary depending on experience, performance and location. Employees in this role will not be paid below the salary threshold for exempt employees in the state where they reside.

Who you are

  • 8+ years of experience in incident management, SRE, technical operations, or a related field, including 2+ years of people management experience (or equivalent combination of education and experience)
  • Proven track record building or scaling an incident management and/or problem management program
  • Strong understanding of structured incident management practices (triage, escalation, post-incident review) and formal problem management methodologies
  • Experience leveraging AI tools or workflows (e.g., LLM-based summarization, automation, analytics) to improve operational speed and output quality
  • Exceptional written and verbal communication skills, with the ability to distill complex technical situations into clear, actionable updates for executives and internal partners
  • Demonstrated sound judgment, composure, and decision-making under pressure during high-severity or ambiguous situations
  • Experience coaching and developing incident managers or similar operational talent
  • Comfortable partnering across Engineering, Product, Support, and business leadership to drive alignment and accountability
  • Plus: Experience with incident management/on-call tooling (e.g., Rootly, JIRA, ServiceNow) and AI-driven operations tooling
  • Plus: Exposure to distributed systems, cloud infrastructure, or large-scale consumer applications

Similar Jobs

More Jobs at Zillow, Inc.

More Information Technology Jobs

Find similar Senior Manager, Incident Management jobs: