Technology Consultant - Site Reliability Engineer SRE

Compunnel

$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years in Site Reliability Engineering, DevOps, or Production Engineering/Support.
  • 4+ years hands-on with Kubernetes, Docker, and containerized environments.
  • 4+ years experience with Java, Spring Boot, microservices, and REST APIs.
  • 3+ years with observability and monitoring tools (e.g., Splunk, Dynatrace, Prometheus).
  • Strong understanding of SLI, SLO, SLA, and SRE principles.
  • Strong analytical, troubleshooting, communication, and problem-solving skills.

Responsibilities

  • Manage and support business-critical applications on Kubernetes and container platforms.
  • Monitor application and platform health to identify performance issues proactively.
  • Troubleshoot Kubernetes deployments, pods, services, and networking issues.
  • Enhance observability solutions for metrics, logs, traces, and alerting.
  • Support Java, Spring Boot, and microservices-based applications.
  • Perform root cause analysis for production incidents and implement corrective actions.
  • Automate operational activities to reduce operational TOIL.

Benefits

  • Opportunity to work with cutting-edge technologies in cloud and containerization.
  • Collaborative environment with cross-functional engineering teams.
  • Focus on reliability engineering and operational efficiency improvement.
  • Professional development and skill enhancement opportunities.
Full Job Description
Job Summary

The Technology Consultant - Site Reliability Engineer (SRE) will provide hands-on expertise in Kubernetes, observability, Java, and production reliability for highly available and distributed enterprise applications. The role will focus on application and platform support, complex production troubleshooting, automation, incident management, and reliability improvements. The engineer will collaborate closely with application engineering, DevOps, cloud, infrastructure, and support teams to improve application availability, scalability, performance, resilience, and operational efficiency.

Key Responsibilities
• Manage and support business-critical applications running on Kubernetes and containerized platforms.
• Monitor application and platform health and proactively identify reliability, availability, and performance issues.
• Troubleshoot Kubernetes deployments, pods, services, networking, configurations, and application issues.
• Implement and enhance observability solutions covering metrics, logs, traces, dashboards, and alerting.
• Support and troubleshoot Java, Spring Boot, and microservices-based applications.
• Perform root cause analysis (RCA) for critical production incidents and implement permanent corrective actions.
• Define and monitor SLIs, SLOs, SLAs, Error Budgets, and other reliability metrics.
• Automate repetitive operational activities and identify opportunities to reduce operational TOIL.
• Participate in incident, problem, change, and production release management activities.
• Collaborate with engineering teams to improve application resilience, performance, scalability, and fault tolerance.
• Support CI/CD pipelines and improve application deployment and release processes.
• Participate in capacity planning, performance tuning, disaster recovery, and production readiness reviews.
• Develop and maintain operational runbooks, troubleshooting procedures, and technical documentation.

Required Qualifications
• 6+ years of experience in Site Reliability Engineering, DevOps, or Production Engineering/Support.
• 4+ years of hands-on experience with Kubernetes, Docker, and containerized application environments.
• 4+ years of experience with Java, Spring Boot, microservices, and REST APIs.
• 3+ years of experience with observability and monitoring tools such as Splunk, Dynatrace, Prometheus, Grafana, Datadog, or ELK.
• Strong understanding of SLI, SLO, SLA, Error Budgeting, and SRE principles.
• Strong analytical, troubleshooting, communication, and problem-solving skills.

Preferred Qualifications
• Experience with Kubernetes deployment and troubleshooting tools such as Helm.
• Experience with AWS, Azure, or Google Cloud Platform.
• Knowledge of Linux/Unix and Shell scripting.
• Experience with Kafka, IBM MQ, or other messaging technologies.
• Knowledge of Terraform, Ansible, or other Infrastructure as Code tools.
• Experience with Jenkins, GitLab CI, GitHub Actions, or Azure DevOps.
• Experience implementing distributed tracing and application performance monitoring.
• Knowledge of incident management and ITIL processes.
• Experience supporting high-volume, highly available, distributed enterprise applications.

Similar Jobs

More Jobs at Compunnel

More Information Technology Jobs

Find similar Technology Consultant - Site Reliability Engineer SRE jobs: