Role descriptionJob SummaryWe are seeking an experienced
ML Ops Engineer to design, build, and support scalable, secure, and production-ready machine learning platforms across cloud and on-premises environments. The ideal candidate will have strong expertise in MLOps, Kubernetes, cloud platforms, automation, and reliability engineering.
Required Qualifications- 8+ years of experience in Platform Engineering, DevOps, MLOps, or related fields.
- Strong experience with GCP (Google Cloud Platform) and cloud-native technologies.
- Hands-on expertise in Kubernetes, including GKE and/or OpenShift.
- Strong proficiency in Python for automation and platform development.
- Experience building and managing MLOps platforms and ML lifecycle workflows.
- Expertise in CI/CD pipelines and infrastructure automation.
- Knowledge of security, data protection, and compliance best practices.
- Experience with observability, monitoring, logging, and incident management.
- Strong understanding of Site Reliability Engineering (SRE) principles.
- Excellent communication and stakeholder management skills.
Preferred Skills- Experience designing enterprise-scale ML platform architectures.
- Multi-cloud experience (AWS, Azure, and GCP).
- Experience supporting AI/GenAI workloads in production environments.
- Knowledge of Infrastructure as Code (Terraform, Ansible, etc.).
- Familiarity with model serving, feature stores, and model monitoring.
- Experience mentoring engineers and driving platform engineering best practices.
- Background in highly regulated enterprise environments.