Perficient

Recovery and Incident Manager with AI-Ops

Perficient$73K — $170K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, IT, Engineering, or related field (or equivalent experience).
  • 8+ years of experience in IT Operations, SRE, Infrastructure Operations, or Production Support.
  • 5+ years leading incident management or reliability engineering efforts.
  • Strong knowledge of SRE, ITSM, ITIL frameworks, and incident management processes.
  • Hands-on experience with monitoring platforms like Dynatrace and ServiceNow for service management.
  • Experience in cloud platforms: AWS, Azure, GCP, and familiarity with DevOps practices.
  • Certifications in ITIL, SRE, and leading cloud platforms.

Responsibilities

  • Monitor and analyze major incident response efforts and service recovery activities.
  • Act as a senior escalation point for operational incidents at Tier 1 and Tier 2.
  • Conduct incident reviews and root cause analysis for corrective action planning.
  • Enhance Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
  • Implement SRE practices for platform reliability and scalability.
  • Define and track service-level agreements (SLAs), objectives (SLOs), and KPIs for operations.
  • Lead AIOps initiatives for automating incident detection and remediation.

Benefits

  • Opportunities for continuous learning and professional development.
  • Flexible work arrangements and locations (Charlotte, Tempe, NYC).
  • Collaboration with cross-functional teams to drive innovation.
  • Access to advanced AI technologies and tools for operational improvements.
Full Job Description
Job Description

We currently have a career opportunity for a Recovery and Incident Manager with AI-Ops to join our team in Charlotte, NC or Tempe, AZ or NYC, NY.

Job Overview:

We are seeking a Recovery and Incident Manager with AI-Ops to lead incident management, operational resilience, and intelligent automation initiatives across enterprise technology environments. This role will partner with Network Operations Center (NOC), Infrastructure Operations, Cloud Engineering, DevOps, and Application Support teams to proactively detect, respond to, and prevent technology incidents. The ideal candidate combines hands-on incident management expertise with experience implementing observability, automation, and AI-driven operational solutions to improve system reliability, reduce operational overhead, and enhance customer experience. The candidate should possess deep expertise in AIOps, ITSM, ITIL, SRE, Incident Management, Cloud Operations, and Enterprise Infrastructure.

Responsibilities

Incident & Recovery Management
  • Monitor, document, and analyze major incident response efforts and service recovery activities.
  • Serve as a senior escalation point for Tier 1 and Tier 2 operational incidents.
  • Conduct incident reviews, root cause analysis, and corrective action planning.
  • Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).

Site Reliability Engineering
  • Implement SRE practices to improve platform reliability, scalability, and resiliency.
  • Define and monitor SLAs, SLOs, and operational KPIs.
  • Develop proactive reliability and availability strategies.

AIOps & Automation
  • Implement AIOps solutions to automate incident detection, diagnosis, remediation, and prevention.
  • Build and optimize AI-powered operational agents and self-healing workflows.
  • Reduce operational effort through intelligent automation.

Observability & Monitoring
  • Lead enterprise monitoring initiatives using Dynatrace and related observability platforms.
  • Improve visibility across cloud, infrastructure, applications, and user experiences.
  • Enable predictive monitoring and anomaly detection.

ITSM & Service Operations
  • Develop and enhance incident, problem, change, and event management frameworks aligned with ITIL and ITSM best practices.
  • Leverage ServiceNow workflow automation to improve service delivery.

Cross-Functional Leadership
  • Partner with Infrastructure, DevOps, Cloud, Security, Application Development, and NOC teams.
  • Mentor operational teams and promote an automation-first culture.


Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field (or equivalent experience).
  • 8+ years of experience in IT Operations, Site Reliability Engineering, Infrastructure Operations, Network Operations, or Production Support environments.
  • 5+ years of experience leading incident management, operational transformation, or reliability engineering initiatives.
  • Strong experience with:
    • Site Reliability Engineering (SRE)
    • IT Service Management (ITSM)
    • ITIL Framework
    • Incident, Problem, Change, and Event Management
    • Network Operations Center (NOC)
    • Infrastructure Operations
    • Service Desk Operations
    • Application Production Support
    • Cloud Platforms (AWS, Azure, or GCP)
    • DevOps Practices and Toolchains
  • Hands-on experience with Dynatrace, monitoring platforms, and observability solutions.
  • Experience using ServiceNow for ticketing, workflow automation, and service management.
  • Strong understanding of infrastructure, networking, cloud architecture, and enterprise application ecosystems.
  • Proven experience conducting root cause analysis and implementing preventive controls.
  • Experience leading enterprise AIOps implementations.
  • Experience building AI-powered operational agents and intelligent automation solutions.
  • Certifications such as:
    • ITIL Foundation or ITIL Managing Professional
    • Certified Site Reliability Engineer (SRE)
    • AWS, Azure, or Google Cloud certifications
    • ServiceNow certifications
  • Experience with workflow orchestration and enterprise automation platforms.
  • Familiarity with predictive analytics, machine learning operations, and autonomous operations frameworks.

ABOUT THE TEAM

Our Automation team empowers organizations to work smarter by connecting digital process automation (DPA), robotic process automation (RPA), and AI into seamless, intelligent workflows. We help leading brands streamline operations, enhance efficiency, and unlock new business potential. By embedding advanced AI models into automation strategies, we enable smarter decision-making, adaptive processes, and continuous optimization at scale.

ADDITIONAL INFORMATION

Applications will be accepted until the position is filled or the posting is removed.

The salary range for this position takes into consideration a variety of factors, including but not limited to skill sets, level of experience, applicable office location, training, licensure and certifications, and other business and organizational needs. The new hire salary range displays the minimum and maximum salary targets for this position across all US locations, and the range has not been adjusted for any specific state differentials. It is not typical for a candidate to be hired at or near the top of the range for their role, and compensation decisions are dependent on the unique facts and circumstances regarding each candidate. A reasonable estimate of the current salary range for this position is $ 73,008 to $ 170,640. Please note that the salary range posted reflects the base salary only and does not include benefits or any potential variable compensation programs. Information regarding the benefits available for this position are in our benefits overview.

#LI-MG1#

About Perficient

Perficient is a leading digital consultancy that helps companies transform their businesses and operations through technology. They deliver solutions to clients that range from Fortune 500 companies to emerging businesses. Perficient has a broad range of capabilities, including strategy, design, technology, and operations. They have expertise in a variety of industries, including healthcare, financial services, retail, and energy. Perficient has been recognized as a top employer and a top company for women technologists. They are committed to giving back to their communities through philanthropy and volunteerism.
Learn more about Perficient
Size
6,079 employees
Market Cap
$2.4 billion
Industry
Net Income
$30.1 million
Founded
1998
5 Year Trend
+9.3%
Revenue
$612.1 million
NASDAQ

Similar Jobs

More Jobs at Perficient

More Information Technology Jobs

Find similar Recovery and Incident Manager with AI-Ops jobs: