Tata Consultancy Services

SRE Operations Lead - AWS, GitLab & AIOps

Tata Consultancy Services$70K — $120K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in Site Reliability Engineering (SRE) or Application Reliability Engineering (ARE)
  • Hands-on expertise in AWS services, including CloudWatch, S3, and Lambda
  • Proficiency with observability tools like Dynatrace, Grafana, and Splunk
  • Extensive experience in incident management, particularly P1/P2 situations
  • Strong programming skills in Python and experience with GitLab CI/CD
  • Knowledge of AIOps, automation strategies, and event correlation
  • Experience with control management in enterprise batch operations using Control-M

Responsibilities

  • Lead and coordinate 24x7 SRE operations for major incident response and recovery
  • Define and enhance service reliability metrics (SLIs, SLOs, SLAs) and improve Mean Time to Repair (MTTR)
  • Implement comprehensive observability using tools like Dynatrace and Grafana
  • Drive automation initiatives including self-healing systems and AI-powered operations
  • Manage AWS platform operations, including complex batch processes and Control-M environments
  • Integrate Claude AI on AWS Bedrock with GitLab for enhanced workflows and automation
  • Develop and maintain Python-based integrations and REST API solutions for GitLab

Benefits

  • Opportunity to work in a cutting-edge cloud environment using AWS and AI technologies
  • Chance to lead critical incident management while improving service reliability
  • Access to advanced observability tools for monitoring and performance improvement
  • Collaboration with geographically distributed teams, offering diverse perspectives
  • Potential for mentorship and career growth in an automation-focused environment
Full Job Description
Must Have Technical/Functional Skills
• SRE / Application Reliability Engineering (ARE) and 24x7 production operations
• Major Incident Management (P1/P2), ServiceNow, and Incident / Problem / Change Management
• SLI, SLO, SLA governance; MTTR reduction; service reliability KPIs
• AWS: CloudWatch, Route 53, S3, CloudFront, Lambda, ECR, and Bedrock
• Observability: Dynatrace, Grafana, Splunk, and CloudWatch
• Control-M and enterprise batch operations
• GitLab CI/CD, REST APIs, Personal Access Tokens, security scanning, and DevSecOps
• Python GitLab library and API-based automation
• Automation, AIOps, event correlation, self-healing, and stakeholder management

Roles & Responsibilities
• Lead 24x7 SRE operations and coordinate P1/P2 major incident response through service restoration and follow-up.
• Own reliability measures including SLIs, SLOs, SLAs, MTTR improvement, and service reliability KPIs.
• Drive end-to-end observability using Dynatrace, Grafana, CloudWatch, and Splunk.
• Lead automation, AIOps, self-healing, event correlation, and AI-driven operations initiatives.
• Oversee AWS platform operations, batch processing, and Control-M environments.
• Integrate Claude AI on AWS Bedrock with GitLab using APIs, PATs, and custom workflows.
• Develop AI-driven analysis of GitLab project data, vulnerabilities, pipelines, and security findings.
• Design GitLab API automation, custom workflows, DevSecOps controls, and CI/CD pipeline improvements.
• Build and maintain Python-based GitLab integrations and REST API solutions.
• Support vulnerability remediation and onboarding/configuration of security scanning tools.
• Manage AWS Lambda, ECR, and Bedrock for deployment and automation; optimize Lambda configuration, concurrency, and scaling.
• Design and support resilient multi-region AWS architectures and containerized deployments using Docker and Amazon ECR.

Generic Managerial Skills
• Lead geographically distributed operations teams and coordinate effectively during critical incidents.
• Communicate reliability risks, service performance, and remediation plans to technical and business stakeholders.
• Drive governance, prioritization, continuous improvement, and cross-team collaboration.
• Mentor engineers and promote automation-first, blameless, and reliability-focused ways of working.

Salary - $70,000 - $120,000 per annum.

About Tata Consultancy Services

Tata Consultancy Services (TCS) is an Indian multinational information technology (IT) services and consulting company, headquartered in Mumbai, Maharashtra, India. It is a subsidiary of Tata Group and operates in 149 locations across 46 countries. TCS is the largest Indian company by market capitalization and is ranked 11th on the Forbes Global 2000 list of the world's biggest public companies. TCS is also the second-largest IT services company in the world by revenue and the largest employer of women in India. The company provides services in areas including IT, consulting, and business solutions.
Learn more about Tata Consultancy Services
Size
469,261 employees
Industry

Similar Jobs

More Jobs at Tata Consultancy Services

More Information Technology Jobs

Find similar SRE Operations Lead - AWS, GitLab & AIOps jobs: