Job Description:
Seeking an experienced Project Manager to lead Digital Site Reliability Engineering (SRE), production support, and cloud operations for business-critical retail, pharmacy, healthcare, and commerce platforms. The role requires strong delivery governance, technical understanding, incident leadership, stakeholder management, and the ability to drive service reliability, operational resilience, automation, and continuous improvement across cloud and on-premises environments.
Key Responsibilities
• Lead end-to-end delivery and governance for Digital SRE, production support, and reliability initiatives across critical applications.
• Oversee service health monitoring, incident, problem, and change management, including major incident command, root-cause analysis, and remediation tracking.
• Define, monitor, and report SLAs, SLIs, SLOs, error budgets, operational risks, dependencies, and service performance.
• Drive operational readiness, capacity planning, performance optimization, certificate management, infrastructure readiness, and production issue prevention.
• Use observability platforms such as Dynatrace, Grafana, Kibana/ELK, and Azure Monitor to support proactive monitoring and performance analysis.
• Promote automation through scripts, runbooks, auto-remediation, CI/CD pipelines, and workflow improvements that reduce manual effort and operational risk.
• Coordinate with business stakeholders, engineering, application, infrastructure, cloud, security, vendor, and offshore teams to resolve issues and deliver agreed outcomes.
• Manage project plans, scope, schedules, resources, risks, escalations, governance forums, status reporting, and knowledge-management activities.
• Mentor technical leads, SREs, developers, and support engineers while strengthening operational standards and service ownership.
Required Skills
• years of overall IT experience, including substantial experience in project management, production operations, application support, or SRE leadership.
• Strong knowledge of Java/J2EE, Spring Boot, microservices, REST APIs, Unix/Linux, Python or shell scripting, SQL, Oracle, and MongoDB.
• Hands-on experience with Microsoft Azure and cloud-native technologies, including Kubernetes and containerized workloads.
• Experience with monitoring and observability frameworks such as Dynatrace, Grafana, Kibana/ELK, and Azure Monitor.
• Knowledge of CI/CD and DevOps tooling such as Jenkins, Azure DevOps, GitHub Actions, Maven, Gradle, or equivalent platforms.
• Demonstrated ability to lead major incidents, facilitate root-cause analysis, manage escalations, and coordinate rapid service recovery.
• Strong project planning, delivery tracking, stakeholder communication, Agile leadership, conflict resolution, negotiation, and team-mentoring skills.
• Experience in digital retail, pharmacy, healthcare, commerce, or other high-availability customer-facing platforms.
Preferred Qualifications
• PMP, PRINCE2, or equivalent project-management certification.
• ITIL v4 or equivalent IT service-management certification.
• Experience preparing RFP responses, business proposals, statements of work, executive reports, and service-review presentations.
• Exposure to Azure Cosmos DB, Postman, application performance monitoring, and enterprise lifecycle or quality-management tools.
Key Competencies
• Leadership and ownership in high-pressure, business-critical environments.
• Strong analytical, troubleshooting, decision-making, and risk-management capabilities.
• Clear executive communication and effective collaboration across technical and business teams.
• Commitment to reliability, automation, continuous improvement, knowledge transfer, and measurable service outcomes.
Work Environment: The position may involve coordination across onsite and offshore teams, participation in production support and major-incident activities, and collaboration with client and partner stakeholders across multiple time zones.
TCS does not use artificial intelligence tools for candidate screening or evaluation. This post is for a current vacancy. The hiring process includes an initial screening, followed by a technical evaluation and managerial discussion.