HPC Infrastructure & Cluster Engineer

D2 Technical Services

$170K — $180K *
Aerospace & Defense
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in Linux systems administration and infrastructure management with a focus on high-performance computing
  • Expertise in managing bare-metal servers and advanced network configurations, including InfiniBand
  • Strong proficiency with workload managers and AI orchestration tools like Run:AI and SLURM
  • Hands-on experience with OpenShift or Kubernetes for container orchestration
  • Experience writing automation scripts in Bash or Python
  • Proven ability to troubleshoot complex hardware, network, and OS-level issues

Responsibilities

  • Manage daily operations of the customer compute cluster, focusing on Linux OS, hardware monitoring, and upgrades
  • Configure and optimize workload management platforms to efficiently handle AI/ML workloads using Run:AI
  • Tune cluster performance at hardware, OS, and network levels for maximum efficiency
  • Administer storage solutions and manage high-speed networking, specifically InfiniBand for low latency
  • Provision and configure environments and container platforms for seamless model deployment
  • Ensure compliance with federal security standards by implementing access controls and maintaining accreditations

Benefits

  • Health/Dental/Vision coverage
  • 401(k) with matching contributions
  • Accrued PTO
  • Short/Long Term Disability and Life Insurance
  • Referral bonuses
  • Professional development reimbursement and more
Full Job Description
**ACTIVE TS/SCI SECURITY CLEARANCE REQUIRED**

We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environment. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Key Responsibilities:
  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations

Basic Qualifications
  • 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand)
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g. Run:AI, SLURM)
  • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes
  • Experience writing automation and configuration scripts (e.g. Bash, Python) to streamline cluster maintenance
  • Proven ability to diagnose and resolve complex hardware, network, and OS-level issues

Preferred Qualifications
  • Familiarity with parallel file systems and high-throughput storage architecture
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies

Additional Information
  • Compensation is unique to each candidate and relative to the skills and experience they bring to the position. The salary range for this position is typically $170-$180k. This does not guarantee a specific salary as compensation is based upon multiple factors such as education, experience, certifications, and other requirements, and may fall outside of the above-stated range.
  • Highlights of our benefits include Health/Dental/Vision, 401(k) match, Accrued PTO, STD/LTD/Life Insurance, Referral Bonuses, professional development reimbursement, and more!

Similar Jobs

More Jobs at D2 Technical Services

More Aerospace & Defense Jobs

Find similar HPC Infrastructure & Cluster Engineer jobs: