Role summary You own the agentic resolution platform end to end - both the cloud-native substrate it runs on and the intelligent agents that run on top of it. From Kubernetes, CI/CD, and observability through agent runtime, LLM patterns, retrieval, evaluation, and human-in-the-loop boundaries, you write the production code, design the architecture, and set the engineering bar that lets automation deflect routine operational work before it ever reaches a human.
This is a senior individual-contributor role that is both strategic and deeply hands-on. You ship production platform services and agentic capabilities, make the build-vs-buy calls on tooling and model access, and partner with the senior architect on the patterns that scale. Success is measured by rising deflection, sustained toil reduction, platform availability and developer velocity, and the ability to scale agentic intelligence through systems and people - not individual effort.
What you'll be doing 1) Agentic resolution platform - intake to action - Architect the agentic resolution platform end to end: intake, classification, action execution, verification, and human-in-the-loop fallback.
- Build the agent runtime and orchestration layer: agent state and memory, tool integration, multi-agent coordination patterns, and confidence-thresholded handoffs to humans.
- Define agent-decision observability, audit-ready posture, and the data contracts that let the platform consume durable fixes from upstream engineering teams.
2) LLM application patterns, knowledge, and evaluation rigor - Design production LLM patterns: prompt engineering, retrieval-augmented generation (RAG), structured outputs, multi-model routing, and hybrid retrieval over the knowledge corpus.
- Own the knowledge-base strategy as a compounding deflection lever - every resolved incident becomes training data and structured retrieval input for future automation.
- Establish evaluation and guardrail frameworks for non-deterministic systems: automated evals, quality scoring, drift detection, and feedback loops that compound agent quality over time.
3) Cloud-native platform - build and operate - Architect and operate Kubernetes (EKS or equivalent) at scale for container and serverless workloads supporting agentic and LLM inference traffic patterns.
- Write production platform services and internal tooling (Python or Go) that automate provisioning, deployment, and operational workflows - not just infrastructure configuration.
- Define and maintain infrastructure as code (Terraform) integrated with a major cloud's AI stack (AWS Bedrock/SageMaker, Azure AI Foundry, or Vertex AI), with secrets management and audit-ready posture for regulated environments.
4) CI/CD, observability, and developer experience - Build and maintain CI/CD pipelines tuned for agentic and AI workloads: model and agent versioning, canary rollouts, evaluation gates, and rollback.
- Own the observability stack (Prometheus/Grafana/OpenTelemetry plus enterprise tooling) and instrument platform health, agent-decision telemetry, and model-inference metrics.
- Establish SLOs, SLIs, and reliability standards for both platform and agentic system health; design self-service patterns and golden-path templates that accelerate delivery.
5) Cross-team partnership, security, and talent development - Partner with the Reliability Engineering team on which production patterns become agent-assisted automations, and with senior architects on the patterns that scale.
- Own security posture: network policies, pod security, secrets rotation, vulnerability scanning, and access controls in a regulated pharmaceutical environment.
- Set the engineering bar through code quality standards, architectural reviews, and role modeling; mentor senior engineers in agentic AI, LLM application patterns, and platform engineering. Influence engineering leaders to adopt automation-friendly patterns at the source, not just downstream.
How You Will Succeed:- Be recognized as the senior technical authority for agentic platform engineering in your area.
- Demonstrate measurable, sustained improvements: rising deflection rates, reduced toil, fewer recurring incidents, faster resolution, platform uptime, and deployment velocity.
- Ship production platform services and agentic capabilities that tangibly move the deflection and developer-experience numbers.
- Scale agentic intelligence through systems, standards, and people - not individual heroics.
Your Basic Qualifications:- Bachelor's degree in Computer Science, Information Technology, or a related technical engineering discipline, including Software Engineering, Computer Engineering, Information Systems, Cybersecurity, Information Science, Network Engineering, Systems Engineering, Computer Information Systems (CIS), Management Information Systems (MIS), Cloud Computing, Data Science
- 7+ years of progressive technology experience with substantial hands-on architecture and delivery of automation, AIOps, agentic systems, or cloud-native platforms at enterprise scale.
- Demonstrated experience designing, building, and deploying agentic AI solutions in production environments, including agent runtimes, multi-agent orchestration, tool integration, and memory/state management, with hands-on implementation using at least one major framework (LangGraph, LangChain, LlamaIndex, or MCP).
- Production experience with LLM application patterns: prompt engineering, retrieval-augmented generation (RAG), structured outputs, and multi-model routing.
- Demonstrated experience designing, deploying, and supporting cloud-native solutions in production environments, including Kubernetes (EKS or equivalent), containerized workloads, infrastructure as code using Terraform, and implementation of AI services on at least one major cloud platform (AWS Bedrock/SageMaker, Azure AI Foundry, or Vertex AI).
- Demonstrated experience designing, developing, and deploying production platform services and engineering tools, with hands-on software development expertise in Python (including asynchronous programming, packaging, and performance optimization) and the ability to deliver scalable, maintainable solutions beyond infrastructure configuration. Experience with Go is a plus.
- Strong CI/CD engineering experience for AI and agentic workloads: model and agent deployment, versioning, canary/blue-green strategies, evaluation gates, and rollback.
- Demonstrated experience implementing evaluation and guardrail frameworks for AI and agentic systems, including automated evaluations, drift detection, human-in-the-loop verification, and feedback loops in production environments.
- Demonstrated experience implementing security controls for cloud-native or production environments, including network policies, secrets management, vulnerability scanning, and access controls within regulated environments.
What You Should Bring:- Working knowledge of operational ML: anomaly detection, event correlation, alert-noise reduction, and model/agent monitoring in production.
- Experience with AI observability tooling (e.g., LangSmith, Weights & Biases, custom telemetry pipelines).
- Hands-on experience with ServiceNow integration patterns and chat-based intake platforms (e.g., Microsoft Teams).
- Experience with vector databases and hybrid retrieval architectures.
- Experience building internal developer platforms (IDP), self-service provisioning, or platform-as-a-product tooling.
- Backend development skills beyond infrastructure tooling: REST/gRPC API design, database design, message queues, distributed systems patterns.
- Demonstrated ownership of measurable deflection or operational-toil-reduction outcomes in a production environment - with the numbers to show it.
- Experience with policy-as-code, automated compliance checks, or shift-left security tooling.
- Observability for AI systems: metrics, logs, traces, and agent-decision telemetry using Prometheus/Grafana/OpenTelemetry, plus at least one enterprise stack (Splunk, Datadog, or Dynatrace).
- AWS cloud platform fluency, including EKS, Bedrock, SageMaker, and cost-optimization tooling.
- Familiarity with FinOps practices and cloud cost governance.
- Experience operating in highly regulated industries (life sciences, financial services, healthcare).
- Prior experience standing up a new engineering capability from a small founding team.
- Frontend experience (React, TypeScript) for internal dashboards and developer portals.
Leadership expectations :
- Acts with enterprise-first mindset, beyond individual products or teams.
- Drives accountability, clarity, and engineering rigor across the team.
- Builds trust through consistency, technical depth, and follow-through.
- Raises the capability of the organization, not just personal output.
- Leads through what they build and how they write, not through org-chart authority.
Additional information Availability to work flexible work hours is/may be required. This team supports continuous operations and may require non-standard work hours, including some work on weekends and holidays.
Actual compensation will depend on a candidate's education, experience, skills, and geographic location. The anticipated wage for this position is
$129,000 - $231,000
Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly's compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.
#WeAreLilly