Manager Incident & Problem Management

Dycom Industries, Inc.

$100K — $120K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in IT, Computer Science, or related field (or equivalent experience).
  • 5+ years in Major Incident Management or IT Service Management in an SAP environment.
  • Deep knowledge of SAP S/4HANA systems and architecture.
  • Expert in root cause analysis (RCA) and maintaining a Known Error Database (KEDB).
  • Experience in managing vendors and enforcing Service Level Agreements (SLAs).
  • Excellent communication skills for drafting executive-level outage communications.
  • Ability to lead technical responses in high-pressure scenarios.
  • Strong analytical skills for incident trend monitoring.

Responsibilities

  • Lead the Major Incident Management process for all critical SAP disruptions.
  • Facilitate emergency technical meetings and coordinate rapid response across teams.
  • Draft and distribute executive communications during service outages.
  • Conduct post-mortem analyses to identify root causes of major incidents.
  • Collaborate with teams to implement permanent fixes to prevent recurring issues.
  • Maintain the Known Error Database to streamline triage efforts.
  • Oversee daily operations of external Application Management Services (AMS) partners.
  • Ensure adherence to SLAs and manage ticket hand-offs between teams.

Benefits

  • Weekly Paychecks
  • Paid Time Off, Parental Leave, and Holidays
  • Comprehensive insurance packages including medical, dental, and life insurance.
  • 401(k) with company match
  • Stock Purchase Plan
  • Education reimbursement
  • Legal insurance and various discounts including gym memberships and pet insurance.
Full Job Description
West Palm Beach, FL

Workplace Type: Office

Employment Type: Salaried

We are seeking a highly resilient Incident & Problem Management Manager to serve as the operational anchor and "first responder" for the SAP CoE. In this critical leadership role within the Service Delivery pillar, you will be responsible for restoring normal service operations and leading comprehensive Root Cause Analysis (RCA).

  • Weekly Paychecks
  • Paid Time Off, Parental Leave, and Holidays
  • Insurance (including medical, prescription drug, dental, vision, disability, life insurance)
  • 401(k) w/ Company Match
  • Stock Purchase Plan
  • Education Reimbursement
  • Legal Insurance
  • Discounts on gym memberships, pet insurance, and much more!


What you'll do

Critical Incident Response
  • Major Incident Management: Lead the Major Incident Management (MIM) process for all Priority 1 (Critical) and Priority 2 (High) SAP disruptions.
  • Emergency Orchestration: Host and facilitate emergency technical bridges, coordinating rapid response efforts across internal IT infrastructure, SAP Basis, network teams, and external vendors.
  • Executive Communications: Draft and distribute clear, business-centric executive communications during outages, keeping the IT Leadership, Corporate Sponsors, and OpCo Presidents continuously informed of impact and estimated recovery times.

Problem Management & RCA
  • Post-Mortems: Lead the post-mortem analysis for all major incidents to effectively decrypt the root cause of systemic failures.
  • Structural Resolutions: Partner with the SAP development and QC teams to design and deploy permanent structural fixes, actively preventing the recurrence of known errors.
  • Knowledge Management: Maintain the CoE's Known Error Database (KEDB) to accelerate and streamline future triage efforts.


AMS & Vendor Governance

  • Vendor Operations: Oversee the daily operations of external Application Management Services (AMS) partners, functioning as the Tier 1 and Tier 2 support teams.
  • SLA Enforcement: Enforce strict adherence to Service Level Agreements (SLAs), holding external vendors strictly accountable for response times, resolution quality, and ticket backlog reduction.
  • Tier Hand-offs: Ensure seamless ticket hand-offs between the external support desk and the CoE's internal Level 3 engineering teams.


Operational Analytics & Trend Spotting

  • Volume Monitoring: Monitor daily incident volumes across all core business domains, including Finance, Operations, and Supply Chain.
  • Training Feedback Loop: Identify spikes in "How-To" tickets that indicate a failure in user training rather than a system defect, and feed this intelligence back to change managers to proactively update training materials.


What you'll need

  • Bachelor's degree in Information Technology, Computer Science, or a related field (or equivalent experience).
  • Minimum of 5+ years of experience in Major Incident Management, Problem Management, or IT Service Management within an SAP environment.
  • Demonstrated expertise in SAP S/4HANA systems and architecture.
  • Strong proficiency in leading root cause analysis (RCA) and maintaining a Known Error Database (KEDB).
  • Proven experience in vendor management, specifically overseeing Application Management Services (AMS) partners and SLA enforcement.
  • Exceptional communication skills, with the ability to draft executive-level communications during critical outages.
  • Proven ability to lead emergency technical bridges and command response efforts under high-pressure situations.
  • Strong analytical mindset for monitoring incident trends and identifying training gaps.


More Jobs at Dycom Industries, Inc.

More Information Technology Jobs

Find similar Manager Incident & Problem Management jobs: