The Optum Insight Engineering team is seeking a Lead Cloud Reliability Engineer with Site Reliability Engineering (SRE) experience to design, build, and scale modern AI Ops solutions across payer and provider transformation initiatives. This is a hands-on technical leadership role requiring deep involvement in architecture and engineering while leading globally distributed teams. The role focuses on delivering SRE Process automation, Infrastructure as code, AI, and Retrieval Augmented Generation (RAG) solutions with enterprise-grade reliability, security, and responsible AI practices.
You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.
Primary Responsibilities: - Automate Infrastructure & Operations: Develop Infrastructure as Code using Terraform, Python, Cloud Infrastructure while adhering to SRE best practices
- Automate SRE practices; Lead incident response; Ensure high availability, scalability, and performance of cloud environments through automation
- Lead end-to-end design and implementation of Automated Ops solutions from concept through production with an emphasis on responsible AI practices
- Design, code, test, and operate software using Python or Node.js
- Design, develop, and deploy AI-powered solutions to address complex business challenges across enterprise scale using RAG based solutions
- Work closely with development and SRE teams to improve system design, advocate for reliability, and mentor fellow engineers
- Leverage enterprise-approved AI tools to streamline workflows, automate tasks, and drive continuous operational efficiency
You'll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.
Required Qualifications: - 8+ years of overall software engineering and Site Reliability Engineering (SRE) experience in a public cloud environment (GCP, AWS, Azure)
- 3+ years of demonstrated hands-on automation experience with Python and Terraform based development
- 2+ years of experience managing Kubernetes environments (EKS, AKS, GKE)
- Available to work rotating 24x7 primary and secondary on-call shifts
Preferred Qualifications: - Bachelor's degree in Computer Science, Software Engineering, or a related technical field (or 8+ years of equivalent software engineering experience in lieu of degree)
- 2+ years delivering AI/ML or Generative AI solutions in production
- CI/CD (GitHub Actions preferred) & Infrastructure as Code experience with Terraform
- Hands-on experience with cloud infrastructure automation, observability tools, and SRE best practices for AI workloads
- Experience leading globally distributed technical teams
- Proven experience building Retrieval Augmented Generation (RAG) pipelines, Agentic AI, or multi-step AI workflows
- Proven effective communication skills with ability to explain complex technical concepts to diverse stakeholders
*All employees working remotely will be required to adhere to UnitedHealth Group's Telecommuter Policy
Pay is based on several factors including but not limited to local labor markets, education, work experience, certifications, etc. In addition to your salary, we offer benefits such as, a comprehensive benefits package, incentive and recognition programs, equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements). No matter where or when you begin a career with us, you'll find a far-reaching choice of benefits and incentives. The salary for this role will range from $112,700 - $193,200 annually based on full-time employment. We comply with all minimum wage laws as applicable.
Application Deadline: This will be posted for a minimum of 2 business days or until a sufficient candidate pool has been collected. Job posting may come down early due to volume of applicants.