DepartmentPSD Enrico Fermi Institute: Gardner Group
Job SummaryAs the Cyberinfrastructure Engineer at the MANIAC Lab within the Enrico Fermi Institute, you will report directly to Research Assistant Professor Giordon Stark. As part of our team (https://maniaclab.uchicago.edu), you will help build and operate advanced cyberinfrastructure for science collaborations such as the ATLAS experiment at the CERN LHC. These collaborations rely on data-intensive, distributed computing technologies such as HTCondor, Kubernetes, and Ceph.
The Cyberinfrastructure Engineer will help operate the storage, compute, and GPU infrastructure in the University's campus data centers. This infrastructure connects researchers to the national-scale OSG fabric and to emerging AI/ML service platforms. The work combines hands-on Linux systems administration and on-site hardware support with cloud-native operations using Kubernetes, Helm, containers, and GitOps tooling. It also involves distributed storage (Ceph via Rook, ZFS), configuration management (Puppet), and monitoring and alerting (Prometheus, Grafana).
The team is collaborative and partly distributed. We meet weekly over Zoom and coordinate daily in Slack. Team members are expected to take ownership of their work and function independently. The position is jointly supported by the ATLAS Midwest Tier 2 Center (MWT2), focused on production operations, and by IRIS-HEP, focused on building the infrastructure that enables AI agents to drive physics analysis and facility operations for the HL-LHC era.
Responsibilities- Facility operations (MWT2 and the UChicago Analysis Facility, approximately 50%)
- Provides systems administrative services for Linux compute clusters (CPU and GPU), storage systems, and related support servers. These systems support ATLAS production and analysis at the Midwest Tier 2 Center and the Analysis Facility.
- Operates, upgrades, and scales Ceph distributed storage (including Rook-managed Ceph on Kubernetes). The goal is to meet growing HL-LHC capacity and throughput requirements.
- Performs on-site hardware work in campus data centers, including installing, retrofitting, and replacing servers, storage, and GPUs. Manages vendor support cases and spare-parts inventory.
- Maintains monitoring, dashboards, and alerting (e.g., Prometheus, Grafana) for storage, compute, network, and hardware health.
- Applies operating system and service security patches and vulnerability mitigations in coordination with University security requirements.
- Performs network diagnostics, throughput measurement, and analysis for both LAN and WAN. Supports the facility's participation in WLCG Data Challenges and HL-LHC readiness scale tests.
- Participates in the team's operations support rotation and maintains documentation and operational runbooks.
- Agentic research infrastructure (IRIS-HEP, approximately 50%)
- Deploys and operates the infrastructure for agent-driven physics analysis. This includes the Lab's MCP gateway, the MCP servers it brokers access to, and the credential services behind them. Together, these let researchers' AI assistants securely use facility resources such as HTCondor, Kubernetes, Rucio, and ServiceX.
- Supports IRIS-HEP integration challenges and demonstrations of end-to-end agentic analysis workflows. The work spans dataset discovery, batch processing, analysis, and inference run through AI agents.
- Operates GPU platforms for machine learning and AI workloads, including locally hosted large language models and sandboxed agent environments.
- Helps develop and operate AI-assisted operations agents that monitor HTCondor, Kubernetes, Ceph, and related services. These agents propose corrective actions under human review.
- Deploys and operates services using container-based approaches (Docker, Kubernetes, Helm) and GitOps workflows (e.g., Flux). Contributes reusable deployment bundles to the SSL Deployment Factory so other facilities can adopt these services.
- Learns new distributed computing, AI/ML, and infrastructure-as-a-service technologies.
- Maintains complex system and network administration functions. Works with moderate guidance to administer simple systems and assists in the administration of larger systems.
- Installs, configures, and maintains operating system workstations and servers. Performs software installations and upgrades to operating systems and layered software packages. Monitors and tunes the system to achieve optimum performance levels, acquiring higher-level skills in the process.
- Performs other related work as needed.
Minimum QualificationsEducation:Minimum requirements include a college or university degree in related field.
Work Experience:Minimum requirements include knowledge and skills developed through 2-5 years of work experience in a related job discipline.
Certifications:---Preferred QualificationsEducation:- Bachelor's degree in Computer Science, Computer Engineering, Physics or related field.
Experience:- Strong experience managing Linux operating systems.
- Configuration management and build systems for large numbers of computers using tools such as Puppet, Chef and Ansible.
- Experience operating distributed storage at scale, particularly Ceph (including Rook on Kubernetes).
- Hands-on data center hardware experience, including server diagnostics, component replacement, and working with vendor support (e.g., Dell iDRAC/warranty processes).
- Experience with monitoring and alerting tools such as Nagios, Prometheus, Grafana, and Alertmanager.
Technical Knowledge or Skills:- Unix/Linux operating systems administration tools and shell scripts.
- Distributed storage systems such as Ceph; local file systems such as ZFS.
- Git version control, scripting (Bash, Python) and automation.
- Knowledge and expertise in technologies such as TCP/IP and related protocols; networked file systems, including NFS.
- Knowledge or experience with batch scheduling systems such as Slurm or HTCondor.
- Knowledge of container technologies such as Docker, Kubernetes, Helm, OpenShift/OKD, OpenStack.
- Familiarity with identity and access management (e.g., Keycloak, OAuth/OIDC).
Preferred Competencies- Strong oral and written communication skills.
- Initiative and capacity for teamwork and creativity.
- Ability to effectively communicate and collaborate with team members, supervisors, and researchers.
- Ability to manage complex technical details and switch between projects.
- Ability to work independently with minimal supervision, take ownership of issues through resolution, and keep the team informed.
- Comfortable collaborating in a distributed team through weekly Zoom meetings and daily Slack communication.
- Knowledge or experience with Spark, Dask or Ray.
- Knowledge of GitOps and cluster lifecycle tooling such as Flux, Argo CD, or Kubespray.
- Experience with GPU servers, including NVIDIA drivers, CUDA, and GPU scheduling in Kubernetes or HTCondor.
- Familiarity with AI/ML infrastructure, such as model serving, LLM-based tooling, or MCP, is a plus.
Working Conditions- Able to work on-site in University data centers on a regular basis to physically install, move, and replace hardware of up to 50lbs/person.
Additional Documents- Resume/CV (required)
- Cover letter (required)
- Professional Reference Information (preferred)
The University of Chicago uses AI-assisted tools to streamline and augment some recruitment processes; however, AI is not used to make hiring decisions.
When applying, the document(s)
MUST be uploaded via the
My Experience page, in the section titled
Application Documents of the application.
Job FamilyInformation Technology
Role ImpactIndividual Contributor
Scheduled Weekly Hours37.5
Drug Test RequiredNo
Health Screen RequiredNo
Motor Vehicle Record Inquiry RequiredNo
Pay Rate TypeSalary
FLSA StatusExempt
Pay Range$80,000.00 - $90,000.00
The included pay rate or range represents the University's good faith estimate of the possible compensation offer for this role at the time of posting.
Benefits EligibleYes
The University of Chicago offers a wide range of benefits programs and resources for eligible employees, including health, retirement, and paid time off. Information about the benefit offerings can be found in the Benefits Guidebook.