About the teamZillow is strengthening how it manages and learns from production incidents, with greater focus on operational discipline, actionable data, improved tooling, and preventing repeat issues. This role offers the opportunity to directly improve system reliability and reduce customer impact.
The role sits at the intersection of technology, operations, and communication. You will independently manage small- to mid-scale incidents while supporting senior incident managers during complex, high-severity events.
The incident management practice is evolving. You will help improve processes, runbooks, tooling, root-cause analysis, and follow-through-not simply respond to incidents.
About the roleZillow Group is seeking an Incident Manager (IC) to support the reliability and operational excellence of our production systems. This role coordinates incident response to minimize customer impact and improve system resilience through structured execution and clear communication.
At this level, success in this role is anchored in three core areas - Operational Execution: independently driving incidents to resolution through structured process and sound judgment; Communication: delivering timely, accurate updates to diverse stakeholders; and Information Synthesis: gathering, connecting, and applying data to drive RCA and continuous improvement.
ResponsibilitiesOperational Execution- Lead non-trivial production incidents from detection through mitigation
- Facilitate incident calls, defining clear ownership, priorities, and next steps
- Execute established incident management workflows, runbooks, and escalation paths
- Coordinate cross-functional teams for timely resolution and minimal customer impact
- Maintain accurate, real-time documentation of incident timelines and actions
Communication- Deliver clear and concise updates to all technical and non-technical stakeholders
- Ensure stakeholder alignment on status, risks, and next steps
- Document incident summaries and contribute to post-incident reporting
- Proactively escalate risks and communication gaps early
Information Synthesis & Continuous Improvement- Support Root Cause Analysis (RCA) by gathering data and documenting contributing factors
- Identify trends, gaps, and opportunities for improving incident workflows
- Contribute to retrospectives and track follow-up actions to completion
- Apply lessons learned to enhance future incident response
Coordination & Collaboration- Partner with Engineering, Product, and Support teams during incident response
- Support Senior Incident Managers on high-severity or complex incidents
- Facilitate collaboration and maintain alignment across stakeholders
Operational Excellence- Maintain a high standard of reliability and responsiveness in on-call rotations
- Reinforce structured incident management practices
- Contribute to the improvement of runbooks, tooling, and team processes
Scope & Impact- Owns end-to-end execution of small to mid-scale incidents
- Contributes to larger, high-severity incidents with guidance
- Influences improvements within team-level processes and workflows
- Builds foundational skills in incident management, communication, and decision-making
This role has been categorized as a Remote position. "Remote" employees do not have a permanent corporate office workplace and, instead, work from a physical location of their choice, which must be identified to the Company. U.S. employees may live in any of the 50 United States, with limited exceptions.
In California, Connecticut, Maryland, Massachusetts, New Jersey, New York, Washington state, and Washington DC the standard base pay range for this role is $106,600.00 - $170,400.00 annually. This base pay range is specific to these locations and may not be applicable to other locations.In Colorado, Hawaii, Illinois, Maine, Minnesota, Nevada, Ohio, Rhode Island, Vermont, and Virginia the standard base pay range for this role is $101,300.00 - $161,900.00 annually. The base pay range is specific to these locations and may not be applicable to other locations.
In addition to a competitive base salary this position is also eligible for equity awards based on factors such as experience, performance and location. Actual amounts will vary depending on experience, performance and location. Employees in this role will not be paid below the salary threshold for exempt employees in the state where they reside.
Who you are- 5+ years of experience in incident management, SRE, or technical operations (or 3 years with a Master's degree, or equivalent work experience)
- Proven experience coordinating responses to production incidents or operational events
- Exceptional written and verbal communication skills, able to tailor updates for diverse audiences
- Demonstrated sound judgment and prioritization in high-pressure or ambiguous situations
- Familiarity with structured incident management processes (triage, escalation, post-incident review) and RCA practices
- Comfortable collaborating with cross-functional teams (Engineering, Product, Support)
- Demonstrates ownership, reliability, and accountability in operational responsibilities
- Plus: Experience with incident management/on-call tooling (e.g., Rootly, PagerDuty, ServiceNow)
- Plus: Exposure to distributed systems, cloud infrastructure, or large-scale applications