Role Overview:This role is for an experienced Engineer specializing in cloud infrastructure, automation, and operations, with a primary focus on Kubernetes and Google Cloud Platform (GCP). The successful candidate will be instrumental in implementing Infrastructure as Code (IaC), managing robust CI/CD pipelines, and providing critical support for production systems. A significant aspect of the role involves applying AI/ML concepts and AIOps practices to enhance operational efficiency and incident response.
Key Responsibilities: - Manage incidents, provide on-call support, and perform production triage to ensure system stability.
- Implement hands-on automation solutions and develop comprehensive CI/CD pipelines.
- Apply strong understanding of AI/ML concepts and AIOps practices, including model lifecycle management, monitoring, and AI-driven alerting.
- Support and operate ML/AI platforms or pipelines (MLOps) as a preferred responsibility.
- Integrate AI-driven automation into existing monitoring and incident response frameworks (preferred).
Required Skills: - Strong experience with Kubernetes and Google Cloud Platform (GKE).
- Strong experience in Infrastructure as Code (IaC) using Terraform, Helm, and GitHub Actions.
- Proficiency in programming languages such as Python, Ansible, and Node.js.
- Strong experience with the Prometheus and Grafana observability stack.
- Solid understanding of Linux systems and networking fundamentals.
- Proven experience in incident management, on-call support, and production triage.
- Hands-on experience with automation and CI/CD pipelines.
- Strong understanding of AI/ML concepts and AIOps practices (model lifecycle, monitoring, or AI-driven alerting).
Qualifications: - 8-10 years of relevant experience.
Preferred Skills: - Google Cloud Architect Certification.
- Certified Kubernetes Administrator (CKA).
- Experience in Java/J2EE and Spring Boot.
- Experience supporting or operating ML/AI platforms or pipelines (MLOps).
- Exposure to AIOps tools, anomaly detection, or predictive analytics systems.
- Experience with large-scale distributed systems and microservices architecture.
- Experience with GPU-based workloads or ML infrastructure on GCP.
- Knowledge of Kubeflow, Vertex AI, or ML pipelines.
- Experience integrating AI-driven automation into monitoring and incident response.