Role Overview:This role involves designing, operating, and supporting Kubernetes platforms, specifically across bare-metal clusters and Google Kubernetes Engine (GKE). The engineer will be responsible for ensuring high availability, scalability, performance, and reliability of production systems, implementing GitOps-based deployment workflows, and managing cloud infrastructure on Google Cloud Platform (GCP).
Key Responsibilities:- Design, operate, and support Kubernetes platforms across bare-metal clusters and Google Kubernetes Engine (GKE)
- Ensure high availability, scalability, performance, and reliability of production systems
- Implement and manage GitOps-based deployment workflows using tools like Argo CD
- Build, maintain, and optimize CI/CD pipelines using tools such as GitHub Actions, Harness, CircleCI, or equivalent
- Deploy and manage applications using Helm, including canary and progressive delivery strategies
- Implement comprehensive observability using Prometheus, Grafana, Loki, and Tempo
- Proactively monitor systems, troubleshoot incidents, and perform root cause analysis (RCA)
- Partner with development teams to improve service reliability, scalability, and operational maturity
- Provision and manage cloud infrastructure on Google Cloud Platform (GCP)
- Automate infrastructure and platform operations using Infrastructure as Code (IaC) and scripting
- Drive continuous improvements in resilience, automation, and operational efficiency
Required Skills:- Strong hands-on experience with Kubernetes architecture and administration
- Experience managing both bare-metal Kubernetes clusters and Google Kubernetes Engine (GKE)
- Solid understanding of Google Cloud Platform (GCP) services and networking concepts
- Proven experience with GitOps practices and tools such as Argo CD
- Proficiency with CI/CD tools (GitHub Actions, Harness, CircleCI, or similar)
- Practical experience with Helm
- Practical experience with Canary / progressive deployments
- Strong expertise in observability and monitoring: Prometheus, Grafana, Loki, Tempo
- Experience with Terraform for infrastructure provisioning
- Infrastructure as Code (IaC) and automation
- Understanding of modern API technologies such as GraphQL
- Familiarity with API management platforms (Apigee Edge, Apigee X)
- Knowledge of CDN and edge services (e.g., Akamai)
Qualifications:Preferred Skills:- Working knowledge of Java (Spring Boot) and/or Node.js frameworks
- Understanding of microservices architecture and service-to-service communication
- Experience with Ansible or similar configuration management tools
- Exposure to hybrid or multi-cloud environments
- Experience in performance tuning and cost optimization on GCP
- Understanding of Kubernetes and cloud security best practices
- SRE experience aligned with SLIs, SLOs, and error budgets
- Strong analytical and troubleshooting skills
- Ownership mindset with a focus on automation and reliability
- Ability to work effectively in a fast-paced, collaborative environment
- Clear communication skills with cross-functional stakeholders