Senior HPC Applications Engineer

Parallel Works

$130K — $155K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years experience with scientific or AI applications on Linux HPC systems
  • Proficient in building complex software on RHEL and Debian/Ubuntu
  • Experience with Spack, EasyBuild, and Lmod in production environments
  • Strong background in multi-node GPU workloads and environment diagnosis
  • Familiarity with application domains like weather, CFD, or materials science
  • U.S. citizenship and eligibility for Secret clearance, with sponsorship available

Responsibilities

  • Compile and package MPI implementations and compilers for user accessibility
  • Install AI/ML frameworks and ensure multi-node GPU launches function correctly
  • Conduct performance analysis through scaling studies and profiling tools
  • Ensure code portability across new GPU architectures using containers
  • Provide user support by diagnosing and resolving failed job issues
  • Create documentation and training resources for user communities

Benefits

  • Medical, vision, and dental coverage
  • 401(k) with company match
  • Short term disability coverage
  • Generous paid vacation and sick time
Full Job Description
About the role

Parallel Works is hiring a Senior HPC Applications Engineer to own the software stack and the user experience on our platforms. Our users are weather modelers, computational chemists, aerospace engineers, and AI researchers. The role is the escalation point for build failures, jobs that die partway through a multi-node run, and jobs running below expected throughput.

The scope covers both long-lived domain codes and current AI workloads, since customers run both on the same clusters. Users also move between on-premises systems, cloud, and commercial GPU providers, so a large part of the job is making an application behave the same across different compilers, site modules, MPI builds, and filesystems. Expect to spend a good share of the week talking to users.
What you will do
  • Software stack: compile and package MPI implementations (OpenMPI, MPICH, Intel MPI, HPC-X), compilers (GCC, Intel oneAPI, NVHPC), and scientific libraries, delivered through Spack or EasyBuild with Lmod module trees users can navigate.
  • Enable AI and ML workloads: install the frameworks customers ask for, get multi-node GPU launch working, and diagnose what sits below the framework: NCCL and collective behavior, container and driver mismatches, storage throughput, node faults mid-run. Customers drive their own toolchain choices.
  • Performance work: run scaling studies, profile with Nsight, VTune, TAU, HPCToolkit, or Score-P, and hand the finding to the systems team when the fix belongs in the fabric or the filesystem.
  • Portability: get customer codes running on new GPU architectures and new venues, using containers where that beats rebuilding against each site's modules.
  • User support: triage tickets, diagnose failed jobs to a root cause, and close them with a written explanation.
  • Documentation and training: user guides, office hours, and training for user communities, including formal Government training events.

Requirements
  • 10 or more years supporting scientific or AI application users on Linux HPC systems.
  • Building complex software from source on both RHEL family and Debian or Ubuntu systems: compilers, MPI, CMake and autotools, and the dependency problems that come with them.
  • Running Spack or EasyBuild and Lmod in production.
  • Experience with multi-node GPU workloads from the platform side, and the judgment to tell a framework problem from an environment problem.
  • Working with site provided software stacks on on-premises systems as well as cloud images where you control the whole stack.
  • Working knowledge of at least one application domain: weather and climate, computational fluid dynamics, molecular and materials science, or structural analysis.
  • United States citizenship and eligibility for a Secret clearance, since the work reaches export controlled Government environments. An active clearance helps. We sponsor candidates who are eligible but not currently cleared.

You do not need every item on this list. If you have most of it and work well with other people, apply.
Preferred Qualifications
  • A prior user facing role at a Government supercomputing center, national laboratory, or university HPC center.
  • Depth in profiling and debugging tools: a GPU profiler, a CPU profiler, gdb, and MPI tooling.
  • Hands-on distributed training or inference work with PyTorch DDP or FSDP, DeepSpeed, Megatron style frameworks, JAX, vLLM, or TensorRT-LLM. Customers own their toolchains, so this is depth rather than a requirement.
  • Jupyter, remote visualization, or virtual desktop support for research users.

Benefits

Medical, vision, and dental coverage, a 401(k) with company match, short term disability, and generous paid vacation and sick time.

Similar Jobs

More Jobs at Parallel Works

More Information Technology Jobs

Find similar Senior HPC Applications Engineer jobs: