Tata Consultancy Services

SRE Operation Lead

Tata Consultancy Services • $100K — $120K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in software engineering or site reliability engineering
  • Proficient in Python and/or Shell scripting
  • Strong programming skills in Java/Spring Boot and/or .NET
  • Experience with REST APIs and Microservices architecture
  • Solid understanding of SQL and database troubleshooting
  • Familiarity with CI/CD pipelines and Git/source control
  • Experience with Linux/Unix environments

Responsibilities

  • Troubleshoot production issues across various layers including UI, APIs, and databases
  • Analyze logs and application errors to identify root causes
  • Develop permanent fixes for recurring production problems
  • Automate operational tasks to reduce mean time to recovery (MTTR)
  • Rapidly create automation scripts for repetitive SRE activities
  • Build self-healing solutions for known production failures
  • Debug and troubleshoot across frontend, backend, and middleware layers

Benefits

  • Discretionary Annual Incentive
  • Comprehensive Medical Coverage including Dental & Vision
  • Family Support with Maternal & Parental Leaves
  • Insurance Options including Home and Identity Theft Protection
  • Convenience benefits like Commuter Benefits and Training Reimbursement
  • Generous Time Off policy including Vacation and Sick Leave
  • Legal and Financial Assistance options including 401K and Student Loan Refinancing
Full Job Description
Must Have Technical/Functional Skills

  • Strong Python and/or Shell scripting
  • Strong programming background in Java/Spring Boot and/or .NET
  • REST APIs and Microservices
  • SQL and database troubleshooting
  • Git/source control
  • CI/CD pipelines
  • Application and production debugging
  • Log analysis
  • Monitoring/observability tools
  • Linux/Unix
  • Automation and self-healing implementation
  • Strong RCA/problem-solving capability


Good to Have

  • New Relic
  • Splunk/Scalyr
  • Docker/Kubernetes
  • Azure or AWS
  • Jenkins/Azure DevOps
  • ServiceNow
  • API Gateway/API Connect
  • Messaging technologies such as Kafka/RabbitMQ/Service Bus


Roles & Responsibilities

Production Troubleshooting & RCA

  • Troubleshoot production issues across UI, APIs, microservices, application, database and infrastructure layers.
  • Analyze logs, application errors, API failures, latency and service dependencies.
  • Perform code-level debugging to identify root causes rather than limiting investigation to monitoring/log analysis.
  • Develop permanent engineering fixes for recurring production problems.
  • Reduce MTTR through automation and faster root-cause identification.
  • Clients current SRE direction specifically emphasizes improved RCA, impact analysis, anomaly detection and faster incident resolution


Automation & Self-Healing

  • Rapidly develop Python/Shell/PowerShell or application-level automation scripts for repetitive SRE activities.
  • Build self-healing and auto-remediation solutions for known production failure scenarios.
  • Automate recurring operational runbooks and SOPs.
  • Develop automated health checks and post-deployment validation.
  • Identify repetitive manual activities and convert them into zero/minimal-touch automation.
  • Integrate monitoring alerts with automated remediation where applicable.
  • This directly aligns with clients documented SRE focus on self-healing systems, auto-remediation scripts, workflow automation, automated platform health checks and automation of recurring SOPs


Microservices & Full-Stack Engineering

  • Strong ability to understand microservices architecture and service-to-service interactions.
  • Debug REST APIs and distributed application flows.
  • Ability to navigate unfamiliar codebases and quickly understand application behavior.
  • Troubleshoot issues across frontend, backend, middleware and data layers.
  • Understand containerized applications and distributed systems.
  • Make application/code/configuration fixes when required instead of depending entirely on development teams.


TCS Employee Benefits Summary:

  • Discretionary Annual Incentive.
  • Comprehensive Medical Coverage: Medical & Health, Dental & Vision, Disability Planning & Insurance, Pet Insurance Plans.
  • Family Support: Maternal & Parental Leaves.
  • Insurance Options: Aut& Home Insurance, Identity Theft Protection.
  • Convenience & Professional Growth: Commuter Benefits & Certification & Training Reimbursement.
  • Time Off: Vacation, Time Off, Sick Leave & Holidays.
  • Legal & Financial Assistance: Legal Assistance, 401K Plan, Performance Bonus, College Fund, Student Loan Refinancing.


#LI-KR3

Salary Range-$100,000-$120,000 a year

About Tata Consultancy Services

Tata Consultancy Services (TCS) is an Indian multinational information technology (IT) services and consulting company, headquartered in Mumbai, Maharashtra, India. It is a subsidiary of Tata Group and operates in 149 locations across 46 countries. TCS is the largest Indian company by market capitalization and is ranked 11th on the Forbes Global 2000 list of the world's biggest public companies. TCS is also the second-largest IT services company in the world by revenue and the largest employer of women in India. The company provides services in areas including IT, consulting, and business solutions.
Learn more about Tata Consultancy Services
Size
469,261 employees
Industry

Similar Jobs

More Jobs at Tata Consultancy Services

More Information Technology Jobs

Find similar SRE Operation Lead jobs: