Technology Consultant - Site Reliability Engineer (SRE)

NTT Data, Inc.

$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years in Site Reliability Engineering, DevOps, or Production Support.
  • 4+ years hands-on with Kubernetes and Docker environments.
  • 4+ years experience with Java, Spring Boot, Microservices, and REST APIs.
  • 3+ years with observability tools like Splunk or Prometheus.

Responsibilities

  • Manage and support business-critical applications in Kubernetes.
  • Monitor application health and issues proactively.
  • Troubleshoot Kubernetes deployment problems and application issues.
  • Enhance observability solutions for performance metrics and alerts.
  • Conduct root cause analysis for production incidents.
  • Define and monitor key reliability metrics like SLIs and Error Budgets.
  • Automate operational activities to reduce toil.

Benefits

  • Opportunities for professional development and training.
  • Flexible work arrangements and remote options.
  • Collaborative and diverse work environment.
  • Access to advanced technology and tools.
Full Job Description
NTT DATA's Client is currently seeking an experienced Technology Consultant - Site Reliability Engineer (SRE) with strong hands-on expertise in Kubernetes, Observability, Java, and production reliability. The ideal candidate will have experience supporting highly available and distributed enterprise applications, troubleshooting complex production issues, and driving automation and reliability improvements.

The role requires close collaboration with application engineering, DevOps, cloud, infrastructure, and support teams to improve application availability, scalability, performance, and operational efficiency.

Day to Day Job Duties
  • Manage and support business-critical applications running on Kubernetes and containerized platforms.
  • Monitor application and platform health and proactively identify reliability, availability, and performance issues.
  • Troubleshoot Kubernetes deployments, pods, services, networking, configurations, and application issues.
  • Implement and enhance observability solutions covering metrics, logs, traces, dashboards, and alerting.
  • Support and troubleshoot Java/Spring Boot and Microservices-based applications.
  • Perform root cause analysis (RCA) for critical production incidents and implement permanent corrective actions.
  • Define and monitor SLIs, SLOs, SLAs, Error Budgets, and other reliability metrics.
  • Automate repetitive operational activities and identify opportunities to reduce operational TOIL.
  • Participate in incident, problem, change, and production release management activities.
  • Collaborate with engineering teams to improve application resilience, performance, scalability, and fault tolerance.
  • Support CI/CD pipelines and improve application deployment and release processes.
  • Participate in capacity planning, performance tuning, disaster recovery, and production readiness reviews.
  • Develop and maintain operational runbooks, troubleshooting procedures, and technical documentation.
Basic Qualifications
  • 6+ years of experience in Site Reliability Engineering, DevOps, or Production Engineering/Support.
  • 4+ years of hands-on experience with Kubernetes, Docker, and containerized application environments.
  • 4+ years of experience with Java, Spring Boot, Microservices, and REST APIs.
  • 3+ years of experience with observability and monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog, or ELK.
Nice to Have
  • Strong understanding of SLI, SLO, SLA, Error Budgeting, and SRE principles.
  • Experience with Kubernetes deployment and troubleshooting tools such as Helm.
  • Experience with AWS, Azure, or Google Cloud Platform.
  • Knowledge of Linux/Unix and Shell scripting.
  • Experience with Kafka, IBM MQ, or other messaging technologies.
  • Knowledge of Terraform, Ansible, or other Infrastructure as Code tools.
  • Experience with Jenkins, GitLab CI, GitHub Actions, or Azure DevOps.
  • Experience implementing distributed tracing and application performance monitoring.
  • Knowledge of incident management and ITIL processes.
  • Experience supporting high-volume, highly available, distributed enterprise applications.
  • Strong analytical, troubleshooting, communication, and problem-solving skills.
#LI-NorthAmerica

Similar Jobs

More Jobs at NTT Data, Inc.

More Information Technology Jobs

Find similar Technology Consultant - Site Reliability Engineer (SRE) jobs: