Site Reliability Engineer

Compunnel

$95K — $115K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience operating Kubernetes in production environments
  • Strong understanding of Linux and command line skills
  • Proven ability to debug and troubleshoot complex systems
  • Experience with public cloud platforms, preferably Azure or AWS
  • Familiarity with automation and observable tooling is a plus

Responsibilities

  • Operate and maintain the Kubernetes-based platform across various cloud environments
  • Utilize observability tools for actionable alerts and maintain updated documentation
  • Manage incidents through the complete lifecycle from response to reviews
  • Onboard and support new clients integrating onto the platform
  • Build automation tools to reduce manual tasks and assess performance
  • Implement process enhancements to optimize operations
  • Manage hardware upgrades and ensure system components are current
  • Monitor capacity proactively to align with performance demands

Benefits

  • Collaborative work environment focused on developer experience
  • Exposure to a wide range of architectures including cloud and hybrid solutions
  • Opportunity to enhance automation and tooling capabilities
  • Engagement with cutting-edge observability tools and practices
  • Support for ongoing professional development and learning opportunities
Full Job Description
JOB SUMMARY
Client API Platform team champions an API first approach across the Firm's development teams, with a consistent focus on delivering an outstanding Developer Experience. We deploy and support solutions across a range of architectures, including cloud, hybrid cloud and on-premises environments. Operational Engineers ensure the platform remains stable, performant and well supported. The role is as much about engineering as it is operations: beyond keeping the platform healthy, it involves building automation and tooling, designing diagnostic tests and engineering practical improvements to how the platform runs. The focus is squarely on running and improving the platform day to day.

Key Responsibilities
Operate and maintain the Kubernetes based platform across public and private cloud environments.
Work with observability tooling to ensure alerts are actionable through up-to-date runbooks and documentation.
Manage incidents from end-to-end, including incident response, root cause analysis, and post incident reviews.
Onboard and support new clients onto the platform.
Build automation and diagnostic tooling that cuts manual effort and evaluates performance.
Identify and deliver process improvements.
Support software and hardware upgrades and keep components up to date.
Proactively manage capacity so the platform scales with demand.

Required Qualifications
Hands on experience operating Kubernetes in production, ideally with Service Mesh.
Strong Linux and command line fundamentals.
Confident debugging & troubleshooting complex systems, from the application layer through to lower-level infrastructure.
Experience working with a public cloud provider, preferable Azure or AWS.

Preferred Qualifications
Working knowledge of Grafana, Prometheus, Loki and Tempo is a plus.
Scripting or coding in Python or Java is a strong plus.
CI/CD, infrastructure as code such as Helm or Terraform is a plus

Similar Jobs

More Jobs at Compunnel

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: