HPC Linux System Administrator

MRI Technologies

$110K — $130K *
Aerospace & Defense
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of Linux system administration experience in a cluster or research computing environment.
  • Hands-on experience administering HPC job schedulers like Slurm, PBS/Torque, or LSF.
  • Production experience with high-speed parallel filesystems such as Lustre or GPFS.
  • Experience with containers in HPC contexts for user environments and workflows.
  • Familiarity with CI/CD workflows tied to HPC clusters and run nodes.
  • Bachelor's degree or equivalent certification in a related field.

Responsibilities

  • Administer the HPC job scheduler, including configuration and troubleshooting job failures.
  • Manage the high-speed parallel filesystem for optimal data handling.
  • Develop and maintain CI/CD workflows for HPC environments.
  • Support scientists and engineers in operational tasks related to human spaceflight.
  • Enhance containerized HPC workflows and assist in system configuration management.

Benefits

  • Comprehensive medical, dental, and vision insurance.
  • Life and disability insurance coverage.
  • Paid time off and a 9/80 work schedule (every other Friday off).
  • 401(k) retirement plan.
  • Opportunity to work within NASA's critical computing environments.
Full Job Description
MRI Technologies, supporting NASA under the JETS II contract, is hiring an HPC Linux System Administrator to help run and improve that cluster - a hands-on, on-premises HPC role administering the job scheduler and parallel filesystem, supporting containerized HPC workflows and CI/CD (including Jacamar), and working directly with the scientists and engineers behind NASA's human spaceflight mission. This is not a public-cloud, DevOps-only, or Kubernetes-only position.

Your HPC experience doesn't have to say "NASA" on it. Another NASA center, a national lab, a university/research computing center, or another government agency all count - and your title doesn't have to have been "System Administrator." We want someone who'll bring fresh ideas from wherever they've been.

Based in Houston, TX at NASA Johnson Space Center - fully remote considered for the right candidate with strong, demonstrated HPC experience.
What We Are Looking For

Required - Core HPC Experience (Must Have)
  • Hands-on, production experience ADMINISTERING an HPC job scheduler (Slurm, PBS/Torque, or LSF) - a hard requirement, not a preference. This means administering the scheduler itself (queue/partition configuration, accounting, troubleshooting job failures), not just submitting jobs to one.
  • Hands-on, production experience ADMINISTERING a high-speed parallel filesystem (Lustre or GPFS) - also a hard requirement. General NAS/SAN/NFS storage administration does not substitute for this.
  • Minimum 5 years of Linux system administration in an on-premises, bare-metal, or cluster/research-computing environment. Cloud-only or DevOps-only backgrounds (e.g., managing EC2/AKS/EKS without on-prem HPC cluster experience) do not meet this requirement.

Required - General
  • Typically requires a bachelor's degree or equivalent certification in a related field, with a minimum of 5 years of experience
  • Experience using containers in an HPC context - packaging and running user environments (including older OS or older package versions) on a shared cluster rather than for traditional microservices
  • Experience building and supporting CI/CD workflows, ideally tied to HPC clusters and run nodes
  • System configuration management experience
  • Experience with monitoring and alerting systems
  • Demonstrated problem-solving, planning, and communication skills
  • Ability to work effectively in a team environment
  • Must be able to provide proof of U.S. Citizenship or U.S. Permanent Residency and complete a U.S. government background investigation (required for this NASA contract - see Important Information below).

Preferences
  • Experience with RedHat-based Linux distributions
  • Familiarity with InfiniBand high-speed networking
  • Experience with provisioning tools (xCAT, Warewulf)
  • Experience with Ansible and/or Foreman for configuration management
  • Familiarity with SPACK software package manager
  • Experience with log consolidation, monitoring, and Git/GitLab (including CI/CD pipelines)
  • Familiarity with Jacamar (https://gitlab.com/ecp-ci/jacamar-ci) or comparable tooling for running CI/CD jobs against HPC clusters and run nodes
  • Experience applying AI in a sysadmin role - integrating AI tooling into operational workflows, automation, and user-facing HPC jobs
  • MPI workflow and administration experience (e.g., Open MPI, MPICH, Intel MPI)
  • Package management and environment modules experience (e.g., Lmod/Environment Modules, SPACK, EasyBuild)
  • Knowledge of NASA security mechanisms (security plans, POAMs, ATOs, Risk Assessments)

This position has been posted at multiple levels. Depending on your experience and business needs, we may consider candidates at any level for which the position is advertised.
Benefits and Perks

We offer a comprehensive benefits package including medical, dental, vision, life and disability insurance, paid time off, and 401(k). You'll also enjoy a 9/80 work schedule (every other Friday off, when applicable),and the chance to work in one of JSC's most critical computing environments supporting human spaceflight.

Similar Jobs

More Jobs at MRI Technologies

More Aerospace & Defense Jobs

Find similar HPC Linux System Administrator jobs: