[8SN] Senior Site Reliability Engineer (SRE) - Kubernetes

Software Mind

$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, or related roles
  • 3+ years of hands-on experience supporting production Kubernetes services
  • Strong production incident response experience and familiarity with runbooks
  • Experience using Splunk for troubleshooting production issues
  • Knowledge of Prometheus and Grafana for building alert rules and dashboards
  • CI/CD and infrastructure-as-code experience for containerized deployments
  • Proficiency in Linux and networking fundamentals, including Kubernetes networking

Responsibilities

  • Support deployment and reliability of production Kubernetes services
  • Monitor health and investigate incidents in distributed applications
  • Engage in on-call support, incident response, and postmortems
  • Troubleshoot application runtime and service issues with engineering teams
  • Facilitate CI/CD and observability for production monitoring
  • Manage client-directed backlogs and established priorities

Benefits

  • Competitive salary and laptop
  • Professional development and training opportunities
  • Work with cutting-edge cloud and container technologies
  • Flexible work arrangements
  • Impact on organization-wide digital transformation initiatives
Full Job Description
Our client is a leading enterprise software company building highly scalable cloud-native platforms used by organizations around the world. Their engineering teams focus on delivering reliable, secure, and high-performing services while embracing modern DevOps, Kubernetes, and cloud technologies. You will join a team responsible for ensuring the stability, reliability, and operational excellence of a critical UI service running in production. Contract Duration: Initial contract through the end of 2026, extending the engagement to a total 12-month term based on performance. About the Role This is a Senior SRE role supporting production reliability for a Kubernetes-based UI service / AI experience framework stack. This is not general infrastructure, and it is not a front-end developer role. The strongest candidates will have production SRE experience across Kubernetes operations, observability, Node.js runtime troubleshooting, JVM / Java service troubleshooting, Splunk, and incident ownership. What You'll Do • Support the deployment, operation, and reliability of production services running on Kubernetes. • Monitor service health and investigate production incidents across distributed applications. • Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements. • Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams. • Support CI/CD, GitOps-based deployments, observability, and production monitoring. • Work within a client-directed backlog and established priorities. Qualifications Required Qualifications • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services. • 3+ years of hands-on production Kubernetes experience strongly preferred. Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting • Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene • Splunk experience for log aggregation, search, and production troubleshooting • Prometheus and Grafana experience, specifically building alert rules and dashboards, not only using existing dashboards • CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux • Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking • Production troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime environment. Experience may include Node.js heap snapshots, CPU profiling, event-loop and memory analysis, as well as JVM GC log analysis, thread dumps, JVM tuning, and Java service latency investigation. • Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication Additional Information Nice to Have • Web Components / Lit experience, to perform first-level debugging of UI-related issues • Server-side rendering or isomorphic runtime experience • Canary rollout / multi-version production operations • Distributed tracing and request-context correlation • KEDA or event-driven autoscaling • Experience with enterprise platform integration layers What We Offer • Competitive salary and laptop • Professional development and training opportunities • Work with cutting-edge cloud and container technologies • Flexible work arrangements and collaborative team environment • Impact on organization-wide digital transformation initiatives

Similar Jobs

More Jobs at Software Mind

More Information Technology Jobs

Find similar [8SN] Senior Site Reliability Engineer (SRE) - Kubernetes jobs: