The ALCF seeks a DevOps engineer to support infrastructure within ALCF and in support of the Department of Energy's American Science Cloud (AmSC) project. AmSC is a secure, federated, and science-optimized cloud environment that integrates the DOE's world-leading computing and experimental facilities, data resources, and high-performance networks.
The AmSC platform enables DOE scientists to create, access, and integrate AI-ready datasets, run scalable AI model training and inference on leadership-class systems, perform distributed simulations, control instruments, and move data efficiently across sites.
As a DevOps Engineer, you will work within the Data Services and Workflows team at ALCF, with significant interaction with the Infrastructure Services group of AmSC to support all activities on our multi-cloud central hub infrastructure, for development, staging, pre-production, and production environments. Other L2 science service teams are deploying services on top of the infrastructure that the Infrastructure team manages - e.g. data catalogs and repositories, at-scale HPC compute services, user interfaces and APIs, and intelligent operations (AI/MLOps). Your primary job responsibilities will be to support the science teams by building foundational infrastructure and developing CI/CD pipelines to deploy services on that infrastructure.
Major Duties/Responsibilities:
- The service stack is primarily Kubernetes-based. Perform cluster administration and application deployment assistance to users.
- Build and maintain pipelines for deploying cloud infrastructure and science services.
- Manage and use image registries such as Harbor.
- Writing and updating automation for resource provisioning and CI/CD pipelines - e.g. Terraform, GitOps, Python.
- Implement security controls as defined by Cybersecurity team (DevSecOps).
- Configure basic instrumentation for infrastructure and core services, to feed into monitoring and alerting systems.
- Provide primary operational support and engineering for production applications.
- Define and implement KPIs, processes and drive continuous improvement.
- Diagnose platform operational problems quickly and effectively.
- Deploy, manage, and operate managed Kubernetes clusters (Amazon EKS, Azure AKS, Google GKE, or equivalent), including node group lifecycle management, cluster upgrades, cloud-native networking integrations (load balancer controllers, CNI plugins), and multi-environment promotion across dev, staging, and production.
- Coordinate with vendors to resolve hardware and software problems.
Hybrid Remote Work - Occasionally Onsite: which applies to employees regularly scheduled for some onsite and some remote days, with employees typically working more than 60% of their time remotely. Non-exempt employees should not be permitted to split their time between on-site and remote work on a given workday unless they have advance supervisor approval.
Position Requirements- PT2: Bachelor's Degree in computer science or closely related field and a minimum of 2+ years of experience as a DevOps engineer and/or Cloud Engineer.
- Previous experience as a team lead - able to perform task management for DevOps or Cloud Engineering teams.
- Excellent interpersonal/communication skills, and the ability to work as part of a team.
- Working knowledge of cloud application architecture patterns and a thorough grasp of common products and managed services for at least one Cloud Service Provider (e.g. AWS).
- Working knowledge of Kubernetes cluster administration and concepts (CR/CRDs) and application deployment strategies (GitOps, Helm).
- Working knowledge of Unix system fundamentals and common network protocols.
- Solid understanding of cloud computing networking concepts.
- Ability to proactively identify performance issues, problems, and areas for improvement.
- Ability to identify requirements and to define, plan, and implement requisite solutions.
- Ability to plan, organize, prioritize tasks, and complete assigned projects with minimal supervision.
- Experience with continuous integration and continuous deployment software methodologies and strategies.
- An understanding of code review and familiarity with tools like GitHub and GitLab.
- Experience using tools such as Nagios, Grafana and Prometheus to monitor systems, metrics, and create dashboards.
- Experience with OpenTofu or Terraform in multi-account AWS environments, including AWS Organizations, SCPs, and IRSA (IAM Roles for Service Accounts).
- Hands-on experience with ArgoCD including App of Apps patterns and ApplicationSets for multi-environment GitOps deployments.
- Familiarity with Tanka, Jsonnet, or equivalent configuration-as-code templating approaches beyond Helm.
- Experience with Kong Gateway or similar API gateway platforms in Kubernetes environments.
- Familiarity with secrets management patterns including AWS Secrets Manager, External Secrets Operator, or comparable Kubernetes-native solutions.
- Ability to model Argonne's Core Values: Impact, Safety, Respect, Integrity, and Teamwork.
- Exposure to high-speed research networks such as ESnet or Internet2 is a plus.
Job FamilyProfessional Technical (PT)
Job ProfileIT Multi-Functional 2
Worker TypeRegular
Time TypeFull time
The expected hiring range for this position is $69,750.00 - $108,810.00.
Please note that the pay range information is a general guideline only. The pay offered to a selected candidate will be determined based on factors such as, but not limited to, the scope and responsibilities of the position, the qualifications of the selected candidate, business considerations, internal equity, and external market pay for comparable jobs. Additionally, comprehensive benefits are part of the total rewards package.
Click here to view Argonne employee benefits!