We are seeking a Lead AI/ML Engineer to lead AI-driven Operations and AIOps capabilities for the Optum Clinical Manager (OCM) Platform-a large-scale, mission-critical healthcare ecosystem supporting Utilization Management, Care Management, Clinical Operations, APIs, cloud services, and integrated applications. In this role, you will drive the vision, design, and deployment of a next-generation Digital Twin for Operations, enabling predictive insights, intelligent automation, incident prevention, automated root cause analysis, operational intelligence, and autonomous remediation capabilities. You will combine deep software/DevOps engineering expertise with technical leadership to deliver AI-powered solutions, mentor teams, and lead operational transformation during major platform incidents and reliability initiatives.
You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.
Primary Responsibilities: - Design, develop, and deploy enterprise-scale AI-powered solutions and Digital Twin capabilities to model platform health, system dependencies, workloads, and business transactions while emphasizing the responsible use of AI
- Leverage enterprise-approved AI tools, machine learning models, Agentic AI, and predictive analytics to identify operational risks, perform failure predictions, and automate decision support before customer impact occurs
- Evaluate emerging trends in AI/ML, AIOps, and observability frameworks to inform solution design, architectural strategy, and continuous operational innovation
- Own and evolve enterprise AIOps capabilities, including AI-driven event correlation, intelligent alert grouping, automated root cause analysis, and predictive change-risk analytics
- Develop AI-powered operational assistants and copilots that accelerate incident triage, diagnostics, resolution, and threat detection
- Lead Major Incident Management (MIM) processes and war-room execution for P1/P2 events across complex cloud, application, infrastructure, API, and network components
- Drive Site Reliability Engineering (SRE) practices across the ecosystem, establishing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and automated workflows to eliminate toil and enhance resiliency
- DevOps, SRE, and platform engineering teams on best practices for AI deployment, automation, monitoring, and operational readiness
You'll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.
Required Qualifications: - 10+ years of experience in Software Engineering, DevOps, Site Reliability Engineering (SRE), Infrastructure Engineering, Cloud Engineering, or Enterprise Operations
- 5+ years of experience leading technical teams or large-scale production operations for critical enterprise platforms
- 3+ years of hands-on experience designing, developing, and deploying AI-powered solutions, machine learning models, or AIOps capabilities in enterprise production environments
- 3+ years of experience with cloud platforms (Azure, AWS, or GCP) and Infrastructure as Code tools (Terraform, ARM, Bicep, or Ansible).
- 3+ years of experience with container orchestration (Kubernetes, Docker, or OpenShift) and CI/CD pipelines (Azure DevOps, GitHub Actions, or Jenkins)
- 3+ years of experience with observability and monitoring platforms (Dynatrace, Splunk, Datadog, Azure Monitor, Grafana, or Prometheus)
- 2+ years of experience leading major incident management (P1/P2) response and war-room execution using incident management platforms (ServiceNow, PagerDuty, xMatters, or JSM)
- 2+ years of experience programming with Python, Java, PowerShell, or Shell Scripting for software development and automation
Preferred Qualifications: - Proven experience with Agentic AI architectures, generative AI applications, large language models (LLMs), or building AI operational assistants and copilots
- Proven experience building Digital Twin platforms or autonomous remediation frameworks for enterprise platforms
- Proven experience mentoring senior engineers, establishing engineering best practices, and influencing architecture across cross-functional teams
- Proven healthcare industry experience with Utilization Management, Care Management, Claims, or Clinical Systems, including knowledge of HIPAA and PHI regulatory controls
*All employees working remotely will be required to adhere to UnitedHealth Group's Telecommuter Policy.
Pay is based on several factors including but not limited to local labor markets, education, work experience, certifications, etc. In addition to your salary, we offer benefits such as, a comprehensive benefits package, incentive and recognition programs, equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements). No matter where or when you begin a career with us, you'll find a far-reaching choice of benefits and incentives. The salary for this role will range from $145,500 - $249,500 annually based on full-time employment. We comply with all minimum wage laws as applicable.
Application Deadline: This will be posted for a minimum of 2 business days or until a sufficient candidate pool has been collected. Job posting may come down early due to volume of applicants.