INFRASTRUCTURE AND CLUSTER ENGINEER (TS/SCI CLEARANCE REQUIRED)

NorthHill Technology

$110K — $130K *
Technical Services
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Active TS/SCI Clearance with ability to obtain CI Polygraph
  • 5+ years in Linux systems administration and infrastructure management
  • Expertise with bare-metal servers, enterprise storage arrays, and network configurations (InfiniBand experience)
  • Strong proficiency in workload managers/jobbing schedulers (Run:AI, SLURM)
  • Hands-on experience with OpenShift or Kubernetes for container orchestration
  • Ability to write automation scripts (Bash, Python) for cluster maintenance
  • Proven troubleshooting skills for complex hardware, network, and OS issues

Responsibilities

  • Manage day-to-day operations of customer compute cluster, including Linux OS administration and hardware monitoring
  • Configure and optimize workload management using Run:AI job scheduler for efficient AI/ML workload distribution
  • Tune cluster performance across hardware, OS, and network to maximize efficiency
  • Administer storage solutions and high-speed networking, transitioning to InfiniBand for low-latency operations
  • Partner with tech integration teams for provisioning environments and dependencies in OpenShift
  • Ensure compliance with federal security standards and maintain system accreditations

Benefits

  • Direct-hire role with a fast-growing Federal Integrator
  • Opportunity for CI Polygraph sponsorship within the first year
  • Engagement in high-performance computing environments
  • Impact on technology integration and performance engineering efforts
  • Work in a critical role within a federal program context
Full Job Description
NorthHill Technology Resources has a need for an Infrastructure and ClusterEngineer to support a Federal Program in Springfield, VA.This is a direct-hire role with our client, a fast-growing Federal Integrator.An active TS/SCI Clearance is required, will sponsor for CI Polygraph within first year.
We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environment under the User Facing and Data Center Services (UDS) contract. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Key Responsibilities:
  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.


Basic Qualifications:
  • Clearance: Active TS/SCI with the ability to obtain CI Poly.
  • Experience: 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
  • Technical Skills:
    • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
    • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
    • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
    • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus: Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.

Preferred Qualifications:
  • Familiarity with parallel file systems and high-throughput storage architectures.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.

Location: Springfield, VA

US Citizenship Required

Similar Jobs

More Jobs at NorthHill Technology

More Technical Services Jobs

Find similar INFRASTRUCTURE AND CLUSTER ENGINEER (TS/SCI CLEARANCE REQUIRED) jobs: