Position Summary:We are seeking a highly skilled and motivated Cloud Senior DevOps Engineer to join our AI Cloud team. In this high-impact role, you will be the backbone of our deployment and infrastructure operations, ensuring that our AI-powered products and platforms are delivered with speed, security, and exceptional reliability. You will act as a crucial bridge between our research/development teams and real-world deployment, driving automation, optimizing cloud-native architectures, and establishing best practices for MLOps and traditional DevOps workflows.
Key Responsibilities:- CI/CD & MLOps Pipeline Management: Design, implement, and maintain end-to-end CI/CD pipelines for both software applications and machine learning models. Automate build, test, deployment, and rollback processes to ensure seamless transitions from innovation to production.
- Cloud-Native & AI Infrastructure: Build, optimize, and scale cloud-native infrastructure using Kubernetes (K8s) and Docker. Manage and provision specialized computing resources (e.g., GPU clusters) to support high-performance AI workloads and model inferencing.
- High Availability Architecture: Take ownership of high-availability design in production environments. Implement disaster recovery (DR) strategies, self-healing mechanisms, capacity planning, and performance tuning to meet stringent business SLAs.
- Infrastructure as Code (IaC): Champion IaC practices utilizing tools such as Terraform, Ansible, and Helm to achieve fully automated, reproducible, and auditable infrastructure provisioning across multiple cloud environments.
- Observability & Monitoring: Architect and refine comprehensive monitoring, logging, and alerting systems (e.g., Prometheus, Grafana, ELK/EFK stack) to provide deep visibility into system health, application performance, and AI model metrics.
- Cross-functional Collaboration: Work closely with R&D, Data Science, Security, and Business teams to streamline workflows, eliminate bottlenecks, and continuously elevate engineering efficiency through Internal Developer Platforms (IDP) and Platform Engineering initiatives.
- Governance, Security & Compliance: Establish and enforce robust system stability and security standards. Manage release workflows, implement Zero Trust access controls, oversee secrets management, and ensure compliance with industry frameworks (e.g., SOC2, ISO27001).
- Incident Management & Resolution: Act as the technical lead during complex system anomalies and major incidents. Spearhead rapid troubleshooting, conduct thorough root cause analysis (RCA), and implement preventative remediation plans.
Basic Qualifications:- Experience & Education: Bachelor's degree or above in Computer Science, Engineering, or a related technical field, with 5+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles.
- Networking & OS: Expert-level knowledge of Linux operating systems and core networking principles (TCP/IP, DNS, HTTP, Load Balancing, VPCs).
- Containerization & Orchestration: Deep mastery of Docker and Kubernetes orchestration, including a thorough understanding of underlying principles, cluster management, and production-level best practices.
- Cloud Platforms: Proven proficiency in designing and managing infrastructure on major Public or Hybrid Cloud platforms (e.g., AWS, GCP, Azure, Alibaba Cloud), including multi-cloud and hybrid-cloud strategies.
- Programming Skills: Strong coding and scripting capabilities in at least one major language (Go, Python, Shell, etc.) with a solid engineering-oriented mindset focused on automation and tooling development.
- Domain Knowledge: Systematic and practical understanding of CI/CD methodologies, Infrastructure as Code (IaC), Observability paradigms, and Site Reliability Engineering (SRE) principles.
- Soft Skills: Exceptional problem-solving abilities, sharp technical judgment, and excellent cross-team communication skills to effectively collaborate in a fast-paced, dynamic environment.
Preferred Qualifications (Plus):- AI/ML Infrastructure Experience: Familiarity with MLOps practices, model serving/inferencing frameworks (e.g., vLLM, TGI, Triton Inference Server), and experience managing GPU clusters for AI/ML workloads.
- Large-Scale Systems: Proven track record working with large-scale distributed systems or high-concurrency environments (e.g., Fintech, Trading, Real-time processing, or AI platforms).
- Platform Engineering: Hands-on experience in designing and building Internal Developer Platforms (IDP) to enhance developer autonomy and productivity.
- Advanced Security: Deep familiarity with Zero Trust architecture, automated security testing (DevSecOps), and implementing strict compliance frameworks (e.g., SOC2, ISO27001).
- Leadership: Prior experience acting as a Technical Lead, mentoring junior engineers, or managing DevOps teams.