The Aerospace Corporation

High-Performance Computing (HPC) Engineer

The Aerospace Corporation$135K — $202K *
Aerospace & Defense
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or equivalent experience
  • 7+ years experience in Linux system administration within an enterprise HPC environment
  • In-depth knowledge of HPC systems, networking, and Linux
  • Proven experience managing the Slurm scheduler for HPC systems
  • Experience with AI & NVIDIA GPU technologies, including CUDA
  • Proficiency in scripting and automation tools like Clush
  • CompTIA Security+ CE certification or equivalent IAT Level II certification

Responsibilities

  • Collaborate with scientists and engineers on mission-critical technical analyses
  • Lead cross-functional teams and provide mentorship to junior engineers
  • Design and implement HPC solutions to optimize resource utilization
  • Manage large-scale classified and unclassified HPC clusters
  • Ensure high-quality infrastructure design and system configuration
  • Develop automation solutions using Infrastructure-as-Code and GitOps practices
  • Monitor and optimize HPC system performance and resource allocation

Benefits

  • Comprehensive health care and wellness plans
  • Paid holidays, sick time, and vacation
  • Telework options and flexible work schedules
  • 401(k) plan with company contributions and immediate vesting
  • Education assistance for career advancement
  • Professional growth and development programs
  • Relocation assistance
Full Job Description

The Aerospace Corporation is seeking a talented High-Performance Computing (HPC) Engineer (Site Reliability Engineer Staff III/IV) to join our Computational Services team. In this role, you will develop, implement, and optimize HPC clusters that support both on-premises and cloud environments. You will work alongside rocket scientists and engineers, tackling complex space enterprise challenges while having direct impact on critical national security missions. We value a collaborative, proactive mindset and a shared commitment to engineering excellence. 

The selected candidate will be required to work full-time, on-site at our facility in El Segundo, CA or Chantilly, VA.

What You’ll Be Doing

  • Collaborate with scientists and engineers on diverse projects supporting mission-critical technical analysis for national space assets 
  • Lead cross-functional teams and mentor junior engineers.  
  • Design and implement HPC solutions that optimize resource utilization across diverse workloads in both classified and unclassified settings.  
  • Manage on premise 10,000-core classified cluster and a 5,000-core unclassified cluster to ensure peak performance.  
  • Deliver high-quality HPC infrastructure design, and system configuration.  
  • Develop and deploy automation solutions using tools such as Clush.  
  • Manage infrastructure using Infrastructure-as-Code and GitOps practices 
  • Implement, support, and optimize GPU computing.  
  • Monitor, analyze, and tune HPC system performance, utilization, and resource allocation to maintain operational efficiency.  
  • Develop cost-efficient HPC service offerings that align with mission and business objectives.  
  • Harden Linux systems to meet stringent security requirements   

What You Need to be Successful

Minimum Requirements for the Site Reliability Engineer Staff III:

  • Bachelor’s degree in Computer Science, Engineering, or equivalent experience.  
  • Minimum of 7 years’ experience in Linux system administration within an enterprise HPC environment.  
  • Experience supporting technical software (compilers, mod&sim tools, languages, COTS, GOTs) including the development of environment modules.  
  • In-depth knowledge of Linux, networking, and HPC systems. 
  • Experience with Infrastructure-as-Code and GitOps  
  • Proven experience in managing the Slurm scheduler and setting up HPC systems for both interactive and batch workloads.  
  • Experience provisioning and supporting AI & NVIDIA GPU technologies (e.g. CUDA) 
  • Proficiency in scripting and competence with automation tools such as Clush.  
  • Experience hardening Linux systems to meet security requirements  
  • Experience with hardware and infrastructure automation in environments using server vendors such as HPE or Cisco.  
  • Strong communication skills, with an ability to work both independently and as part of a geographically distributed team.  
  • CompTIA Security+ CE certification or equivalent that meets DoD 8570.01-m requirements for IAT Level II personnel  
  • Ability to obtain and maintain a TS/SCI clearance (U.S. citizenship required). 
  • Demonstrated ability to lead cross-functional teams and mentor junior engineers.  

In addition to the above, the minimum requirements for the Site Reliability Engineer Staff IV include:

  • 9+ years of experience in an enterprise 100+ server HPC cluster operations and administration 
  • Experience with performance analysis and optimization with custom developed technical software in collaboration with scientists and engineers. 
  • Expertise in optimizing and customizing Slurm partitions, qualities of service and priority to balance utilization and reduce job wait times.  
  • Experience performing in-place upgrades of Slurm.  
  • Implementing visualization of live system telemetry 
  • Experience developing and architecting cluster configuration management  
  • Advanced Infrastructure-as-Code GitOps (e.g. multi-branch pipelines) 
  • Skill in provisioning and supporting AI & NVIDIA GPU technologies (e.g. CUDA, Nsight), with expertise in GPU integration, resource allocation, and scheduling using Slurm.  

How You Can Stand Out

It would be impressive if you have one or more of these:

  • An active TS/SCI clearance with CI Polygraph. 
  • Experience integrating Slurm with SELinux 
  • Experience implementing DISA STIG compliance 
  • Experience integrating Kubernetes and Slurm, e.g. Slinky 
  • Experience supporting and managing diverse HPC workloads, including computational fluid dynamics, Monte Carlo, structural analysis 
  • Experience integrating user web portals to launch & manage workloads (e.g. Open OnDemand and developing plugins for session persistence and VSCode) 
  • Experience implementing utilization dashboards (e.g. XDMod) with Slurm 
  • Experience managing parallel file systems such as Lustre.  
  • Experience developing solutions that optimize data storage  
  • Experience implementing or supporting Slurm REST API 
  • Experience with automated provisioning, e.g. Warewolf, Kickstart, PXE 
  • Knowledge of NVLINK and DCGM for optimizing GPU workflows.  
  • Familiarity with Prometheus and Grafana for monitoring and performance visualization.  
  • Background in containerization within an HPC context used for data processing and technical analysis  
  • Experience packaging custom software (e.g. RPMs) 
  • Proficiency with automation tools such as Ansible for HPC  
  • Hands-on background with cloud HPC services  
  • Experience with AWS Parallel Computing Service (AWS ParallelCluster).  

We offer a competitive compensation package where you’ll be rewarded based on your performance and recognized for the value you bring to our business.  The grade-based pay range for this job is listed below.  Individual salaries within that range are determined through a wide variety of factors including but not limited to education, experience, knowledge and skills.

(Min - Max)

$135,200.00 - $202,800.00

Pay Basis: Annual

Leadership Competencies

Our leadership philosophy is simple: every employee, regardless of level and role, can demonstrate leadership. At Aerospace, our commitment is our people. To cultivate our talent and ensure that we have a strong pipeline of future leaders, we want individuals who:

  • Operate Strategically
  • Lead Change   
  • Engage with Impact   
  • Foster Innovation   
  • Deliver Results  

Ways We Reward Our Employees

During your interview process, our team will provide details of our industry-leading benefits.

Benefits vary and are applicable based on Job Type.  A few highlights include:

  • Comprehensive health care and wellness plans

  • Paid holidays, sick time, and vacation

  • Standard and alternate work schedules, including telework options

  • 401(k) Plan — Employees receive a total company-paid benefit of 8%, 10%, or 12% of eligible compensation based on years of service and matching contributions; employees are immediately eligible and vested in the plan upon hire

  • Flexible spending accounts

  • Variable pay program for exceptional contributions

  • Relocation assistance

  • Professional growth and development programs to help advance your career

  • Education assistance programs

  • An inclusive work environment built on teamwork, flexibility, and respect

We are all unique, from various backgrounds and all walks of life, yet one thing bonds all of us to each other—the belief that we can make a difference. This core belief empowers us to do our best work at The Aerospace Corporation.

About The Aerospace Corporation

The Aerospace Corporation is a nonprofit corporation that operates a federally funded research and development center (FFRDC) headquartered in El Segundo, California. The corporation provides technical guidance and advice on all aspects of space missions to military, civil, and commercial customers. Aerospace also designs and develops spacecraft, sensors, and other systems in support of national security, civil, and commercial customers. The corporation has more than 4,000 employees and operates a number of laboratories and test facilities across the United States.
Learn more about The Aerospace Corporation
Size
4,000 employees
Industry
Founded
1960

Similar Jobs

More Jobs at The Aerospace Corporation

More Aerospace & Defense Jobs

Find similar High-Performance Computing (HPC) Engineer jobs: