Provide continuous eyes-on-glass monitoring of digital platforms, telemetry, customer-signal indicators, change activity, infrastructure dashboards, and application health signals to support early detection and rapid triage.
Correlate alerts, alarms, incidents, PIMS activity, high-volume days, DML requests, banner requests, and partner escalations to determine business impact, urgency, ownership, and next-best action.
Coordinate escalation and communication across Digital Incident Response, Messaging, Product Service Management, 3L engineering, development, Consumer Technology, Cyber Security, TMIM, and other responder groups.
Support major incident response by moving quickly from signal detection to triage, escalation, executive visibility, stakeholder updates, emergency banner readiness, and coordinated stabilization actions.
Prepare and maintain accurate incident communications using standardized structures that clearly identify event status, customer or business impact, actions underway, accountable owners, and next update timing.
Support customer and stakeholder communication modernization, including subscription-based notification models, emergency banner fast-path processes, and reliable communication pathways that reduce noise and missed stakeholders.
Contribute to governance improvements for traffic routing, single-intake workflows, AMA-orchestrated execution, role-based approvals, evidence capture, paging-group remediation, and repeatable operational controls.
Participate in structured handoffs, health checks, runbook execution, training, onboarding, and knowledge-sharing routines that strengthen follow-the-sun coverage and global team consistency.
Support AI-assisted incident response pilots by helping define use cases, evaluate alert enrichment, draft summaries, validate escalation recommendations, and measure improvements in triage speed and decision velocity.
Track operational outcomes and support reporting for metrics such as mean time to detect, mean time to acknowledge, mean time to restore, incidents identified before broad customer impact, banner activity, PIMS support, traffic-routing requests, DML volume, broadcasts, and ad hoc communications.
Strong attention to detail, accuracy, operational discipline, and ability to maintain composure during high-pressure incident scenarios.
Excellent verbal, written, and interpersonal communication skills, with the ability to translate technical signals into clear business-impact narratives for leaders and stakeholders.
Experience supporting incident response, major incident coordination, command-center operations, production monitoring, technology service management, or digital operations.
Ability to correlate data across multiple monitoring, infrastructure, application, change, customer-signal, and communication sources to identify risk, impact, and escalation paths.
Experience with structured handoffs, runbooks, health checks, stakeholder updates, executive summaries, emergency banner processes, or customer-impact communication workflows.
Strong time management skills, with the ability to multitask, prioritize competing operational needs, and manage multiple requests during active incidents or high-volume change windows.
Experience collaborating across global teams, including U.S. and India shift models, partner responders, engineering teams, product service managers, and business stakeholders.
Working experience with ServiceNow, Splunk, AppDynamics, APM tools, Glassbox, DownDetector, telemetry dashboards, incident notification tools, or comparable monitoring and service-management platforms.
Familiarity with AI-assisted operations, alert enrichment, automated signal correlation, executive summary drafting, or incident-response automation concepts.
Demonstrated ability to improve processes, reduce operational noise, strengthen governance, document procedures, and support measurable improvements in detection, acknowledgement, escalation, and restoration outcomes.
The average range of working hours for this role will be between 10 AM and 9 PM CT.
Ability to work weekends and holidays as assigned.
Ability to work on call as assigned and support a 24x7 follow-the-sun operating model.
Ability to work additional hours as needed during major incidents, high-risk change windows, customer-impacting events, or periods of elevated operational volume.
Flexibility to operate in a fast-paced production-support environment that requires rapid triage, accurate communication, disciplined documentation, and timely escalation.
Ability to maintain readiness for emergency banner support, stakeholder notifications, incident bridge participation, and executive visibility updates when customer or business impact is probable or confirmed.
Reflected is the base pay range offered for this position. Pay may vary depending on factors including but not limited to demonstrated examples of prior performance, skills, experience, or work location. Employees may also be eligible for incentive opportunities.