Job DescriptionManager, Major Incident & Problem Management
Remote Position
Annual Compensation: $95,000 - $108,000 DOE
Overview We are seeking a Manager to own the execution and continued maturity of two critical IT Service Management practices, Major Incident & Problem Management. This role leads the response to all enterprise Priority 1 incidents, coordinates rapid restoration of business services, and ensures leaders and stakeholders receive timely, accurate, and business-focused communications.
Beyond active incident response, this position sets the strategic direction for Major Incident and Problem Management. The role provides functional, dotted-line leadership to the existing MSP (Managed Service Provider) Outage Coordinators and Problem Analyst, establishes consistent operating standards, drives accountability for root-cause and corrective-action work, and uses operational insights to reduce repeat incidents and improve service reliability.
This is a hands-on leadership role for someone who can remain composed during high-impact events, bring structure to ambiguity, influence teams without relying on direct authority, and translate technical conditions into clear business impact and decisions.
What You'll Do
Enterprise Major Incident Leadership
- Own and lead the end-to-end response for all enterprise P1 incidents, from declaration and bridge activation through service restoration, stakeholder transition, and formal closure.
- Establish command and control during major incidents by clarifying roles, driving urgency, maintaining decision discipline, and ensuring the right technical and business resources are engaged.
- Facilitate incident bridges, maintain focus on restoration, remove coordination obstacles, and escalate risks or resource gaps to technology leadership.
- Ensure business impact, scope, workarounds, recovery progress, and restoration status are validated before they are communicated.
- Coordinate executive, technology, and business communications using clear, concise, and audience-appropriate messaging.
- Lead post-incident reviews and confirm that key decisions, timelines, lessons learned, and follow-up actions are documented.
Problem Management Ownership
- Own the enterprise Problem Management practice, including intake, prioritization, investigation governance, known-error discipline, and closure criteria.
- Ensure significant and recurring incidents are evaluated for problem records and that root-cause analysis is completed with appropriate rigor.
- Drive accountable corrective and preventive actions with named owners, target dates, evidence of completion, and risk-based escalation for overdue work.
- Partner with engineering, infrastructure, application, vendor, and service-owner teams to eliminate systemic causes and reduce recurrence.
- Identify patterns across incidents, problems, changes, monitoring events, and service dependencies to inform reliability priorities.
Functional Team Leadership
- Provide dotted-line leadership, operating direction, coaching, and quality oversight for MSP provided Outage Coordinators and the Problem Analyst.
- Define role expectations, coverage models, escalation paths, facilitation standards, documentation requirements, and communication quality controls.
- Conduct case reviews and targeted coaching to build consistency, confidence, and sound judgment across the team.
- Coordinate workload and coverage with internal leaders and vendor management while maintaining clear accountability for practice outcomes.
- Serve as the escalation point for complex incidents, stalled investigations, unresolved ownership, and process exceptions.
Strategy, Governance & Continuous Improvement
- Develop and maintain the multi-year strategy, roadmap, operating model, policies, procedures, playbooks, and maturity plan for Major Incident and Problem Management.
- Establish governance forums and performance reviews that focus on outcomes, risks, recurring failure themes, corrective-action health, and improvement priorities.
- Define and monitor meaningful measures such as restoration performance, communication timeliness and quality, recurrence, root-cause completion, action aging, and business impact.
- Identify opportunities to automate workflows, notifications, evidence capture, reporting, and handoffs across ITSM systems and adjacent platforms.
- Align the practices with IT Service Management standards and integrate them with Change, Configuration, Knowledge, Event, Service Level, and Continuity Management.
- Create training and simulation exercises that strengthen incident leadership, technical response, business-impact assessment, and executive communication.
What You'll Bring
Experience
- 5+ years of progressive experience in IT Service Management, service operations, incident management, problem management, or a related enterprise technology function.
- Demonstrated experience leading high-severity incidents in a complex, multi-team environment with material business impact.
- Experience designing, maturing, or governing Major Incident and Problem Management processes, not only executing individual cases.
- Experience leading internal teams, managed-service providers, or matrixed resources through influence and clearly defined accountability.
- Experience presenting incident status, risk, root cause, and corrective-action progress to senior technology and business leaders.
Operational & Technical Depth
- Strong working knowledge of ITIL practices, particularly Incident Management, Major Incident Management, Problem Management, Change Enablement, Configuration Management, Knowledge Management, and Service Level Management.
- Practical experience with an enterprise ITSM platform; ServiceNow experience is strongly preferred.
- Ability to understand complex application, infrastructure, network, cloud, integration, and vendor dependencies sufficiently to lead restoration and challenge assumptions.
- Ability to use incident and problem data to identify trends, quantify operational risk, and prioritize improvement opportunities.
- Comfort with on-call or after-hours engagement when enterprise P1 incidents require leadership.
Leadership & Communication
- Calm, decisive, and highly organized during fast-moving, high-pressure events.
- Exceptional facilitation skills with the ability to maintain urgency without creating noise or confusion.
- Clear writer and communicator who can translate technical detail into business impact, decisions, risks, and next steps.
- Strong judgment, ownership, follow-through, and willingness to escalate when service restoration or corrective action is at risk.
- Collaborative and credible with technical teams, business stakeholders, executives, and external partners.
Education
Bachelor's degree in Information Technology, Computer Science, Business, or a related field, or equivalent practical experience.
Preferred Certifications
- ITIL 4 or 5 Foundation; ITIL Practice Manager, Monitor, Support and Fulfil, or equivalent advanced ITSM certification.
- ServiceNow Certified System Administrator, Certified Implementation Specialist - IT Service Management, or equivalent platform experience.
- Relevant incident command, problem analysis, reliability, or project leadership certification.
The application window for this position is anticipated to close on 10/31/26.
Learn how our values are at the core of our services and vital to how we approach care and check out our comprehensive benefit options at GlobalMedicalResponse.com/Careers.
More Information about this Job
Check out our careers site benefits page to learn more about our benefit options.