Job DescriptionWe are seeking a highly skilled Lead Site Reliability Engineer (AI & Cloud Operations) to drive the reliability, scalability, automation, and operational excellence of our cloud-native platforms and AI-powered solutions. This role will serve as a technical leader responsible for building resilient infrastructure, implementing modern DevOps and SRE practices, and enabling enterprise AI capabilities through automation, observability, and operational intelligence.
The ideal candidate combines deep expertise in AWS cloud technologies, Kubernetes, infrastructure automation, CI/CD, and incident management with hands-on experience supporting AI/ML and Generative AI platforms. This individual will partner closely with software engineering, data engineering, machine learning, security, and product teams to establish highly available systems, streamline deployments, optimize platform performance, and accelerate innovation through AI-driven operations.
Responsibilities- Lead the design, implementation, and continuous improvement of Site Reliability Engineering (SRE) practices to ensure highly available, scalable, and resilient cloud platforms.
- Architect, deploy, and support AWS-based infrastructure and services, including containerized and serverless environments.
- Build, maintain, and optimize CI/CD pipelines and Infrastructure as Code (IaC) solutions to accelerate and standardize deployments.
- Develop automation solutions, operational tooling, and self-healing capabilities using Python, Shell, and modern DevOps technologies.
- Manage Kubernetes and container platforms, ensuring performance, scalability, and operational stability.
- Establish and enhance observability through monitoring, logging, alerting, and performance management tools to proactively identify and resolve issues.
- Lead incident response, root cause analysis, problem management, and service reliability improvement initiatives.
- Partner with engineering, data, AI/ML, and security teams to support enterprise applications, analytics platforms, and cloud-native solutions.
- Design and implement AI-driven operational capabilities, including intelligent monitoring, automated remediation, predictive analytics, and chatbot-enabled support workflows.
- Support MLOps and AI platform operations, including model deployment, monitoring, governance, and lifecycle management.
- Define and track reliability metrics, service-level objectives (SLOs), and operational KPIs to drive continuous improvement.
- Mentor and provide technical leadership to engineering teams while promoting best practices in reliability, automation, cloud operations, and AI-enabled innovation.
Qualifications- 5+ years of experience in Site Reliability Engineering (SRE), DevOps, or Cloud Operations.
- Strong hands-on experience with AWS services including EC2, EKS, ECS, Lambda, S3, RDS, IAM, CloudWatch, and VPC.
- Experience building and managing CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI/CD, Azure DevOps, or similar platforms.
- Strong scripting and automation skills using Python, Shell, or similar languages.
- Experience with Infrastructure as Code (IaC) tools such as Terraform or CloudFormation.
- Expertise in containerization and orchestration technologies (Docker, Kubernetes).
- Experience with observability tools such as Prometheus, Grafana, Datadog, Splunk, or ELK Stack.
- Understanding of analytics platforms, data pipelines, and operational data analysis.
- Strong troubleshooting, problem-solving, and incident management skills.
- AI & Automation Experience
- Experience implementing AI/ML or Generative AI solutions within enterprise environments.
- Familiarity with AI platforms such as Azure OpenAI, AWS Bedrock, Amazon SageMaker, OpenAI APIs, LangChain, or NVIDIA AI ecosystem.
- Experience building AI-assisted operational workflows, chatbots, intelligent monitoring, predictive analytics, or automated remediation solutions.
- Understanding of MLOps concepts, model deployment, monitoring, and governance.
Applications will be accepted until the position is filled or the posting is removed.
The salary range for this position takes into consideration a variety of factors, including but not limited to skill sets, level of experience, applicable office location, training, licensure and certifications, and other business and organizational needs. The new hire salary range displays the minimum and maximum salary targets for this position across all US locations, and the range has not been adjusted for any specific state differentials. It is not typical for a candidate to be hired at or near the top of the range for their role, and compensation decisions are dependent on the unique facts and circumstances regarding each candidate. A reasonable estimate of the current salary range for this position is $111,300 to $144,600. Please note that the salary range posted reflects the base salary only and does not include benefits or any potential variable compensation programs. Information regarding the benefits available for this position are in our benefits overview.
Disclaimer: The above statements are not intended to be a complete statement of job content, rather to act as a guide to the essential functions performed by the employee assigned to this classification. Management retains the discretion to add or change the duties of the position at any time.
#LI-RS1