What You'll Be DoingThe HPC Systems Administrator will build and operate scalable AI and high-performance computing platforms that power advanced machine learning and scientific workloads. You will partner with AI scientists, engineers, and domain experts to enable efficient model training, inference, and experimentation across GPU, cloud, and on-premises environments. Through platform engineering and automation, you will improve productivity, performance, and access to advanced computing resources while ensuring the reliability, availability, and efficiency of Lilly's AI and HPC infrastructure.
How You'll Succeed- Deliver highly available, secure, and performant AI and HPC platforms that meet the needs of research and engineering teams.
- Drive operational excellence through automation, standardization, monitoring, and continuous improvement.
- Balance infrastructure reliability, scalability, and cost efficiency across environments.
- Enable efficient ML workflows through automation for orchestration, resource scheduling, data access, and reproducibility
- Collaborate effectively across scientific, engineering, and infrastructure teams to solve complex technical challenges.
- Adapt quickly to evolving AI, GPU, and HPC technologies and translate new capabilities into business value.
What You Should Bring- Deep expertise in Linux systems administration, automation, and infrastructure management, with strong scripting skills in Python, Bash, and/or Ansible.
- Experience building, administering, and optimizing large-scale HPC, GPU, or AI/ML computing environments, including job scheduling and resource management platforms such as Slurm or Grid Engine.
- Proficiency with automation, configuration management, and container technologies, including Ansible, Kubernetes, Docker, and related tooling.
- Solid understanding of distributed computing, high-performance networking, storage architectures, and cluster infrastructure.
- Experience supporting large-scale distributed training and inference workloads across multi-GPU and multi-node environments.
- Knowledge of GPU infrastructure, hardware lifecycle management, monitoring, and observability practices.
- Demonstrated ability to solve complex infrastructure challenges, identify root causes, and implement scalable automated solutions.
- Good communication and collaboration skills, with the ability to work effectively across researchers, engineers, and infrastructure teams.
- Experience running NVIDIA GPU infrastructure and hardware lifecycle operations, including GPU monitoring, partitioning, diagnostics, and out-of-band server management using industry-standard tools and protocols.
- Experience supporting regulated or critical environments is a plus.
- Experience operating AI/HPC infrastructure in cloud environments such as AWS, Azure, or GCP.
Your Basic Qualifications- Bachelor's Computer Science, Electrical/Computer Engineering, Systems Engineering, or a related technical field.
- 5 years' experience deploying, administering, or supporting large-scale HPC, GPU, or distributed computing environments in an enterprise or research setting.
Location & Work FlexibilityThis role is based at our Silicon Valley Hub. We offer a flexible hybrid work model, with
three days onsite and two days working remotely each week, supporting both collaboration and work-life balance.
Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE[redacted]), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women's Initiative for Leading at Lilly (WILL).
Actual compensation will depend on a candidate's education, experience, skills, and geographic location. The anticipated wage for this position is
$141,000 - $231,000
Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly's compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.
#WeAreLilly