HPC Infrastructure & Cluster Engineer

Abile Group, Inc.

$108K — $130K *
Aerospace & Defense
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • TS/SCI clearance with eligibility for CI Polygraph requirements
  • Bachelor's Degree in a related discipline or equivalent experience
  • 5+ years of Linux systems administration focused on high-performance computing
  • Expertise in managing bare-metal servers, enterprise storage, and InfiniBand networks
  • Strong proficiency with workload managers and AI orchestration tools like Run:AI
  • Experience with OpenShift or Kubernetes for container orchestration
  • Automation scripting skills in Bash or Python

Responsibilities

  • Manage daily operations of the compute cluster, including Linux OS, hardware monitoring, and upgrades
  • Configure and optimize workload management using Run:AI for AI/ML distributions
  • Tune cluster performance to maximize efficiency across hardware, OS, and network levels
  • Administer storage solutions and high-speed networking, including InfiniBand infrastructure
  • Work with integration teams to provision environments using Red Hat OpenShift for model deployment
  • Ensure compliance with federal security standards and maintain system accreditations

Benefits

  • Opportunities for professional development and training
  • Supportive work environment focused on innovation
  • Engagement in a long-term contract within the Intelligence Community
  • Exposure to cutting-edge technologies in a high-performance computing context
  • Chance to contribute to critical national security projects
Full Job Description
Overview

Abile Group has an exciting and challenging opportunity for a HPC Infrastructure & Cluster Engineer on a 10 year contract providing User Facing and Data Center Services supporting an Intelligence Community customer. All the personnel on the team will work together to support innovative design, engineering, procurement, implementation, operations, sustainment and disposal of user facing and data center information technology (IT) services on multiple networks and security domains, at multiple locations worldwide, to support the IC mission.

 

The right candidate will possess the belowskills and qualificationsand be ready to handle all responsibilities independently and professionally.

Responsibilities
  • Cluster Administration:Manages the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management:Configures, maintains, and optimizes workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization:Tunes cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management:Administers storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration:Partners with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance:Ensures all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
Qualifications

Clearance Required: TS/SCI with ability to obtain a CI Poly.

 

Degree and Years of Experience: Bachelor's Degree in a related discipline, or the equivalent combination of education, professional training, or work/military experience.

  • 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.

Required Certifications:

  • Meet DoD 8570 IAT Level II requirements including one of the following:Security+ CE, CND, SSCP, GSEC, GICSP, CySA+, or CCNA.

Required Skills:

  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
  • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus:Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.

 

Desired Skills:

  • Familiarity with parallel file systems and high-throughput storage architectures.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.

 

Similar Jobs

More Jobs at Abile Group, Inc.

  • VTOC Operations
    $80K — $95K *
    Saint Louis, MO 63129 (Saint Louis County)
    Technical Services
    In-Person
  • Cloud Engineer
    $110K — $130K *
    St. Louis, MO 63129 (Saint Louis County)
    Information Technology
    In-Person
  • Systems Administrator (SaaS)
    $95K — $115K *
    Springfield, VA 22153 (Fairfax County)
    Information Technology
    In-Person
  • Sr Network Engineer
    $108K — $130K *
    Springfield, VA 22153 (Fairfax County)
    Telecommunications & Hardware
    In-Person
  • Sr Network Engineer
    $110K — $130K *
    St. Louis, MO 63129 (Saint Louis County)
    Information Technology
    In-Person

More Aerospace & Defense Jobs

Find similar HPC Infrastructure & Cluster Engineer jobs: