Role Overview:This role is for a Site Reliability Engineer (SRE) with a strong emphasis on Google Cloud Platform (GCP) and AI skills. The position requires robust incident management capabilities within the GCP ecosystem, focusing on maintaining high reliability and operational efficiency.
Key Responsibilities:- Proactively reduce noise from monitoring systems to enhance alert accuracy and operational focus.
- Effectively manage and resolve high-priority incidents, ensuring minimal disruption and rapid recovery.
Required Skills:- Strong experience in GCP Cloud environments, complemented by excellent incident management skills.
- Hands-on proficiency with Google Cloud Platform services including Pub/Sub, BigQuery, Spanner, Dataflow, and Firestore.
- Experience utilizing Vertex AI for model training and the deployment of AI pipelines.
- Familiarity and experience working with Agentic AI assistants.
- Strong expertise with containerization and orchestration technologies such as Kubernetes, GKE (Google Kubernetes Engine), Docker, and Helm.
Qualifications:- 6-8 years of relevant professional experience.
Preferred Skills:- Experience with Amazon Web Services (AWS) services, including EC2, S3, and Lambda.