High-Performance Computing (HPC) Systems Engineer

Jefferson Lab

$91K — $145K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of enterprise Linux system administration experience
  • Familiarity with high availability production environments
  • Automation of test procedures experience
  • Bachelor's degree in Computer Science or related field
  • Preferred: Master's degree in Computer Science or Data Science

Responsibilities

  • Design and implement HPC software infrastructure using containerization and orchestration tools
  • Deploy and maintain physical infrastructure for scientific computing services
  • Create highly available services from hardware to user applications
  • Engage users to ensure computing service performance
  • Develop software to improve existing tools
  • Audit and recommend architectural changes for performance and security
  • Contribute to upstream development of community tools

Benefits

  • Opportunities to work on cutting-edge research projects
  • Collaboration with students and interns in research
  • Contribution to community-driven development efforts
  • Focus on high-performance, large-scale computing environments
Full Job Description
The good-faith pay range for this role is $91,800 - $145,050 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate's skills, experience, education, certifications, and work location.

What your job will be like:

We are seeking a High-Performance Computing (HPC) Systems Engineer to architect, deploy, and maintain the large-scale physical hardware, distributed filesystems, and low-latency networking infrastructure powering our scientific computing ecosystem. This role focuses on the bare-metal and system-level foundations of a petabyte-scale environment, ensuring high availability, peak storage performance, and reliable data movement for experimental nuclear physics workloads. The ideal candidate will blend deep Linux systems administration expertise with modern infrastructure-as-code automation to support state-of-the-art research computing clusters.

In this job you will:
  • Collaboratively design and implement software infrastructure supporting of High Performance and High Throughput computing using best-in-class containerization and orchestration tools.
  • Deploy, maintain, and operate physical infrastructure for scientific computing services.
  • Create highly available services through consideration of the entire stack from hardware and networking though the user application.
  • Engage with users to meet service requirements for performance and availability.
  • Develop software to address gaps in existing tools.


Additional Responsibilities
  • Audit and recommend architectural changes to improve performance, security, and availability.
  • Contribute to upstream development efforts of community tools.
  • Work with students and interns on research projects.


Experience
  • Required: 3 or more years Experience performing enterprise linux System Administration tasks including installation, configuration, and support of COTS/GOTS/FOSS software, file, network, and large-scale storage systems.
  • Required: Experience operating in a production environment with high availability requirements.
  • Required: Experience with automating test procedures
  • Preferred: Experience with NP/HEP HPC infrastructure like slurm, Rucio, Globus, XrootD


Education
  • Required: Bachelor's Degree in Computer Science, Computer Engineering, or related degree with significant Computer Science coursework
  • Preferred: Master's Degree Computer Science, Data Science, or related discipline


Experience and Education Exchange

Education above the minimum may be substituted for experience. Relevant experience may not be substituted for education.

Knowledge, Skills, and Abilities
  • Ability to architect, provision, tune, and maintain petabyte-scale, high-performance parallel and distributed storage systems, including Lustre, Ceph, CephFS, and NFS.
  • Expert Kubernetes, Docker/Podman, and containerization administration.
  • Expert Knowledge of Linux system administration including installation and configuration management (ie. Ansible/Puppet/Foreman)
  • Ability to develop, debug, and test applications based upon design and performance requirements.
  • Ability to communicate clearly in writing (e.g., email, presentations, drawings) and explain work to their supervisor and others in the group.
  • Ability to work effectively with peers and participate in participate in troubleshooting, including off-hours during outages and as part of an occasional on-call rotation.


Similar Jobs

More Jobs at Jefferson Lab

More Information Technology Jobs

Find similar High-Performance Computing (HPC) Systems Engineer jobs: