Role Overview:This role involves leading and managing an SRE/Production Support team, with a primary focus on ensuring client expectations and service level objectives are consistently met. The ideal candidate will possess a strong understanding of ITSM and Site Reliability Engineering (SRE) processes to effectively manage production support operations.
Key Responsibilities:- Lead and manage the SRE/Production Support team to ensure client expectations and service level objectives are consistently met.
- Possess a strong understanding of ITSM and Site Reliability Engineering (SRE) processes to effectively manage production support operations.
- Prepare and maintain operational reports and documentation, including bridge updates, RCA tracking, incident trends, service availability, platform health metrics, monthly operational deliverables, and trend analysis.
- Facilitate technical discussions and bridge calls for critical and escalated production issues, ensuring effective communication and stakeholder alignment throughout the incident lifecycle.
- Participate in Incident Management bridge calls, driving timely resolution by coordinating with engineering teams and escalating issues to the appropriate Subject Matter Experts (SMEs) as required.
- Monitor platform health and identify opportunities to improve system reliability, availability, observability, and operational efficiency.
- Drive continuous service improvement initiatives by analyzing recurring incidents, identifying root causes, and recommending preventive actions.
- Plan, coordinate, and facilitate Disaster Recovery (DR) exercises, ensuring readiness, documentation, and post-exercise review of outcomes.
Required Skills:- Leadership and management of SRE/Production Support teams.
- Strong understanding of ITSM and Site Reliability Engineering (SRE) processes.
- Proficiency in preparing and maintaining operational reports and documentation (e.g., bridge updates, RCA tracking, incident trends, service availability, platform health metrics, monthly operational deliverables, trend analysis).
- Ability to facilitate technical discussions and bridge calls for critical and escalated production issues.
- Experience in Incident Management, including coordinating with engineering teams and escalating to SMEs.
- Skills in monitoring platform health and identifying improvements in system reliability, availability, observability, and operational efficiency.
- Capability to drive continuous service improvement initiatives through incident analysis, root cause identification, and preventive action recommendations.
- Experience in planning, coordinating, and facilitating Disaster Recovery (DR) exercises.
Qualifications: