Anticipated End Date:
2026-09-21
Position Title:
Principal Cloud Engineer – AI/ML
Job Description:
Principal Cloud Engineer – AI/ML
Location: This role requires associates to be in-office 1 - 2 days per week, fostering collaboration and connectivity, while providing flexibility to support productivity and work-life balance. This approach combines structured office engagement with the autonomy of virtual work, promoting a dynamic and adaptable workplace. Alternate locations may be considered if candidates reside within a commuting distance from an office.
Please note that per our policy on hybrid/virtual work, candidates not within a reasonable commuting distance from the posting location(s) will not be considered for employment, unless an accommodation is granted as required by law.
PLEASE NOTE: This position is not eligible for current or future visa sponsorship.
At Elevance Health, Cloud Engineering builds and operates secure, scalable platforms that help application teams provision environments, deploy workloads, manage cloud cost, and transform applications. The Principal Cloud Engineer - AI/ML is a hands-on principal engineer within the Cloud organization. This role designs, codes, integrates, and operates AI-enabled capabilities across account vending and landing zones, Containers as a Service (CaaS), FinOps, developer onboarding, and application migration and transformation.
The Principal Cloud Engineer – AI/MLcreates production agents, reusable agent skills, and Model Context Protocol (MCP) servers and integrations that make cloud services easier to consume and operate. The role partners with Enterprise AI, Information Security, Responsible AI, Enterprise Architecture, and application teams to apply enterprise standards and approved AI services within cloud products. A substantial portion of the role is hands-on build and production ownership. Success is measured through working software, customer adoption, faster onboarding and provisioning, improved developer experience, lower cloud cost, more reliable operations, and accelerated application transformation.
How you will make an Impact:
Deploy AI-enabled capabilities for core Cloud Engineering products, with regular hands-on work in application code, infrastructure as code, APIs, CI/CD pipelines, and production environments.
Develop production agents and reusable skills for cloud customer journeys, including account, subscription, and project requests; landing-zone configuration; access and policy validation; deployment guidance; incident triage; CaaS troubleshooting; cost optimization; and migration readiness.
Build and maintain MCP servers, adapters, and supporting services that securely expose cloud platform APIs, automation, inventory, telemetry, cost data, and operational knowledge to approved agents and assistants.
Advance account vending and developer onboarding by automating intake, prerequisite checks, guardrail validation, documentation, environment setup, access workflows, and handoff to platform services, reducing time to first productive deployment.
Enhance CaaS offerings through self-service and AI-assisted capabilities for workload onboarding, deployment, policy compliance, observability, troubleshooting, and guided remediation across Kubernetes-based environments.
Create FinOps tooling and agents for allocation and tagging quality, anomaly detection, rightsizing, forecasting, commitment utilization, waste reduction, and actionable optimization recommendations; measure realized savings and cost avoidance.
Build application transformation accelerators for discovery, dependency analysis, cloud readiness, containerization, migration planning, modernization recommendations, and generation of repeatable implementation artifacts.
Partner with Enterprise AI, Information Security, Responsible AI, and Enterprise Architecture to consume standards, integrate approved models and tooling, complete security and risk reviews, and implement identity, data protection, audit, evaluation, and lifecycle controls.
Instrument capabilities with product and operational metrics, support production incidents, gather cloud-customer feedback, and continuously improve adoption, reliability, latency, cost, and developer experience.
Provide technical leadership through reference implementations, code reviews, pairing, and mentoring while remaining accountable for delivering working components.
Maintains components of architecture strategy and vision.
Maintains enterprise level blueprints.
Coordinates all enterprise-level conceptual architecture components (e.g., data architecture, application architecture, technical architecture.
Monitors usage of architectural components and assumes responsibility for reuse.
Drives system migration based upon roadmaps defined in enterprise and domain blueprints.
Leads architecture strategy and vision for enterprise.
Ensures blueprints are refreshed as needs emerge or in accordance to plan of record changes.
Provides continuous consulting services and direction in projects and architectures.
Champions and responsible for enterprise level technology and architectural standards, guidelines, principles, frameworks, and reference models.
Minimum Requirements:
Requires an BA/BS degree in Information Technology, Computer Science or related field of study and a minimum of 8 years experience in architecture/design in relevant technology disciplines; or any combination of education and experience, which would provide an equivalent background.
Preferred Skills, Experiences and Competencies:
Demonstrated recent hands-on experience building and operating production cloud platforms and software.
Ability to author and review application code, infrastructure as code, APIs, automated tests, CI/CD pipelines, and operational tooling.
Deep experience with at least one major public cloud provider and core cloud platform patterns, including organization and account, subscription, or project provisioning; landing zones; identity and access management; networking; policy; logging; and monitoring.
Strong experience building or operating Kubernetes and container platform capabilities, including workload onboarding, deployment automation, security controls, observability, reliability, and lifecycle management.
Strong software engineering skills in one or more languages such as Python, Go, TypeScript, or Java, with experience building APIs, event-driven integrations, SDKs, automation services, and reusable platform components.
Hands-on experience building production generative AI or agent solutions, including tool or function calling, orchestration, context and retrieval, state management, structured outputs, automated evaluation, observability, and guardrails.
Experience building MCP servers, clients, or comparable agent-to-tool integration capabilities, including authentication and authorization, schema and version management, error handling, policy enforcement, and audit logging.
Experience with infrastructure as code and automation tools such as Terraform or CloudFormation, secure CI/CD and DevSecOps practices, secrets management, least-privilege access, and production support.
Working knowledge of FinOps practices and cloud cost management, plus experience supporting application migration, modernization, containerization, or platform onboarding programs.
Experience integrating agents into developer portals, IDEs, command-line tools, chat interfaces, or workflow automation platforms.
Experience implementing responsible AI controls, AI security testing, agent evaluation, and human approval patterns in production.
Experience working in regulated environments, preferably healthcare, and collaborating effectively across matrixed enterprise teams.
Job Level:
Non-Management Exempt
Workshift:
1st Shift (United States of America)
Job Family:
IFT > IT Architecture