UnitedHealth Group

Principal AI/ML Platform Engineer - Remote

UnitedHealth Group$164K — $282K *
US-AnywhereRemote in Eden Prairie, MN
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree or 4+ years of equivalent software/platform engineering experience in lieu of a degree
  • 10+ years of experience in infrastructure, DevOps, SRE, or ML platform engineering
  • 5+ years of experience operating Kubernetes or OpenShift at scale in production environments
  • 3+ years of experience designing and managing accelerated-compute infrastructure utilizing NVIDIA GPU technologies
  • 3+ years of experience architecting hybrid AI platforms with major cloud services
  • 3+ years of experience with HPC/AI networking technologies
  • 3+ years of experience managing multi-cluster fleets using GitOps tooling

Responsibilities

  • Own the end-to-end reference architecture for multi-tenant AI compute platforms
  • Set network and latency standards for distributed training
  • Establish cluster governance and automated lifecycle management
  • Design cost and utilization models for AI infrastructure
  • Standardize model-serving platforms and inference gateways
  • Define workload placement frameworks for deployment decisions
  • Design and enforce security controls for AI workloads
  • Establish enterprise SLOs and upgrade strategies for platforms

Benefits

  • Comprehensive benefits package
  • Incentive and recognition programs
  • Equity stock purchase option
  • 401k contribution
  • Flexibility to work remotely from anywhere in the U.S.
Full Job Description
As a Principal AI/ML Platform Engineer on the UnitedHealth Group (UHG) enterprise team, you will serve as the AI Architect across IaaS and PaaS environments, owning the technical direction, reference architecture, and long-term evolution of our multi-tenant AI compute platform end to end. Our team builds and maintains an advanced compute estate spanning on-premises bare-metal Red Hat OpenShift AI clusters equipped with high-performance NVIDIA GPUs and InfiniBand/RoCE training fabrics, alongside public-cloud managed AI services including Azure AI Foundry, AWS Bedrock, and GCP Vertex AI. In this role, you will define architecture standards, optimize high-throughput model training and inference pipelines, establish cost and utilization economics, and enforce strict HIPAA, security, and data-governance standards for regulated healthcare workloads.

You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.

Primary Responsibilities:

  • Own the end-to-end reference architecture for multi-tenant AI compute platforms across hybrid on-premises bare-metal OpenShift AI clusters and public-cloud managed AI platforms (Azure AI Foundry, AWS Bedrock, GCP Vertex AI)
  • Set network, latency, and topology standards for distributed training (NVLink, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL), ensuring interconnect boundaries are strictly maintained
  • Establish cluster governance, GitOps workflows (Argo CD), RHACM policies, and automated lifecycle management for bare-metal accelerated compute nodes
  • Design cost and utilization models including capex amortization, accelerator-sharing strategies (MIG/time-slicing), and cost-per-token/training-run math to inform accelerator procurement roadmaps
  • Standardize model-serving platforms (vLLM, KServe) and inference gateways, setting quantization policies, provenance review gates, and intelligent model routing rules
  • Define workload placement frameworks to determine self-hosted versus managed cloud deployment based on data residency, latency, cost, and compliance requirements
  • Design and enforce identity, access, and security controls for AI workloads and autonomous agents, including least-privilege RBAC, Vault secret management, short-lived credentials, and mTLS
  • Establish enterprise Service Level Objectives (SLOs), disaster recovery plans, and upgrade strategies for OpenShift, OpenShift AI, GPU operators, drivers, and firmware
  • Partner with cross-functional AI teams, LLM gateway engineers, privacy, and security stakeholders to ensure seamless integration and HIPAA compliance


You'll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.

Required Qualifications:

  • Bachelor's degree or 4+ years of equivalent software/platform engineering experience in lieu of a degree
  • 10+ years of experience in infrastructure, DevOps, SRE, or ML platform engineering
  • 5+ years of experience operating Kubernetes or OpenShift at scale in production bare-metal or enterprise cloud environments
  • 3+ years of experience designing and managing accelerated-compute (AI/GPU) infrastructure utilizing NVIDIA GPU Operator, NFD, MIG/time-slicing, and DCGM
  • 3+ years of experience architecting hybrid AI platforms spanning self-hosted IaaS and public-cloud PaaS managed AI services (e.g., Azure AI Foundry, AWS Bedrock, or GCP Vertex AI)
  • 3+ years of experience with HPC/AI networking technologies, including InfiniBand or RoCEv2, GPUDirect RDMA, and NCCL collective communication limits
  • 3+ years of experience managing multi-cluster fleets using RHACM (or equivalent) and GitOps tooling (Argo CD or Flux)
  • 3+ years of experience implementing enterprise security and IAM controls for software workloads (RBAC, OIDC/OAuth, Vault secrets management, mTLS)


Preferred Qualifications:

  • Experience with distributed training frameworks (PyTorch DDP/FSDP, DeepSpeed, Ray, JAX) and batch scheduling systems (Kueue, Volcano)
  • Hands-on experience with LLM inference serving technologies (vLLM, TensorRT-LLM, KServe) and platform tooling such as OpenShift AI (RHOAI) or Kubeflow pipelines
  • Experience with high-performance parallel storage systems (Ceph/ODF, Lustre, IBM Storage Scale, VAST, WEKA)
  • Active Red Hat certifications (e.g., Red Hat Certified Architect / RHCA) or open-source contributions to CNCF, OpenShift, or AI infrastructure projects
  • Experience operating AI/ML platforms within regulated healthcare environments under HIPAA and UHG data privacy controls


*All employees working remotely will be required to adhere to UnitedHealth Group's Telecommuter Policy.

Pay is based on several factors including but not limited to local labor markets, education, work experience, certifications, etc. In addition to your salary, we offer benefits such as, a comprehensive benefits package, incentive and recognition programs, equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements). No matter where or when you begin a career with us, you'll find a far-reaching choice of benefits and incentives. The salary for this role will range from $164,600 - $282,200 annually based on full-time employment. We comply with all minimum wage laws as applicable.

Application Deadline: This will be posted for a minimum of 2 business days or until a sufficient candidate pool has been collected. Job posting may come down early due to volume of applicants.

About UnitedHealth Group

UnitedHealth Group is a medical facility. They offer health technology, health finance, and pharmacy services. They are utilizing clinical data and intelligence to assist in the redesign, automation, and deployment of technology to streamline administrative operations and clinical decision-making. Payment systems for both consumers and providers are critical components of a health-care system.

UnitedHealth Group Careers

Joining UnitedHealth Group means becoming part of a diverse team dedicated to making a difference. As a leader in health services and innovation, our company offers a variety of job opportunities that allow professionals to leverage their skills, drive innovation, and improve lives. Work You’ll Do At UnitedHealth Group, you’ll contribute to a mission-focused environment where your expertise will influence the health and well-being of people worldwide. Our employment opportunities span across a wide range of disciplines, from healthcare specialists to data analysts, ensuring that your career journey is both dynamic and rewarding. Transform Healthcare with Your Expertise UnitedHealth Group stands at the forefront of health innovation. Our team collaborates to deliver solutions that lead to better patient outcomes. Working with us, you’ll find yourself at the intersection of technology, healthcare, and leadership, providing key insights that drive industry transformation. Join Our Global Team As part of our team, you’ll engage with over 300,000 professionals globally, dedicated to building a diverse and inclusive workplace. UnitedHealth Group is not just a company; it’s a community where you can grow your career through continuous learning and leadership opportunities. Our commitment to diversity training ensures that every team member can thrive. UnitedHealth Group Career Development We are committed to your professional growth. Explore career paths filled with promising job opportunities and internships that will harness your potential and expand your capabilities. Whether you’re a seasoned professional or a recent graduate, you’ll find that our career development programs support your ambition at every level. Innovative Work Environment At UnitedHealth Group, innovation is at the core of our operations. We encourage our team to bring forward-thinking ideas that challenge the status quo and lead to breakthrough improvements in patient care. Be Part of a Great Team Our culture fosters a collaborative and supportive environment where every member’s contribution is valued. Enjoy the benefits of being part of a global team that’s committed to making a difference in people’s lives. Future-Proof Your Career With UnitedHealth Group, your career is future-proof. Dive into a range of positions that offer both challenges and rewards. Our robust support system includes unmatched training, development programs, and certification support to propel your career forward. Stay Connected Join Our Team Discover the right position that matches your skills and interests. We are always hiring and look for passionate, curious, and solution-driven team players. Search UnitedHealth Group jobs today and take the first step towards a fulfilling career. Keep Up to Date Stay informed with career tips, insider perspectives, and industry-leading insights you can use today—all from the people who work here. Read Careers Blog Job Alert Emails Customize your subscription to receive job alerts, the latest news, and insider tips tailored to your preferences. Explore the exciting and rewarding opportunities that await at UnitedHealth Group.
Learn more about UnitedHealth Group
Size
350,000 employees
Market Cap
$493.1 billion
Industry
Net Income
$15.4 billion
Founded
1974
5 Year Trend
+9.2%
Revenue
$257.1 billion
NASDAQ

Similar Jobs

More Jobs at UnitedHealth Group

More Information Technology Jobs

Find similar Principal AI/ML Platform Engineer - Remote jobs: