HPC Infrastructure & Cluster Engineer

Bridge Core (BCore)

$148K — $179K *
Aerospace & Defense
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Active TS (SCI eligibility) clearance and ability to obtain CI poly
  • Bachelor's degree
  • 5+ years experience in Linux systems administration and infrastructure management for high-performance computing
  • Expertise in bare-metal servers, enterprise storage, and advanced network configurations
  • Proficiency in workload managers and AI orchestration tools such as Run:AI
  • Hands-on experience with OpenShift or Kubernetes for container orchestration
  • Ability to write automation scripts in Bash or Python

Responsibilities

  • Manage the administration and health of the customer compute cluster
  • Configure and optimize workload management using Run:AI job scheduler
  • Tune cluster performance across hardware, OS, and network levels
  • Administer storage solutions and high-speed networking infrastructures
  • Provision and configure environments for customer model deployment using OpenShift
  • Ensure compliance with federal security standards and maintain accreditations

Benefits

  • Recognition programs including service anniversaries and spot awards
  • Opportunity to work with experienced IT engineering professionals
  • Comprehensive health, dental, and vision insurance
  • 401(k) retirement plan with company contributions
  • Paid time off and life insurance options
  • Employee referral bonuses and stipends
Full Job Description
Overview
HPC Infrastructure & Cluster Engineer Springfield, VA

Active TS (SCI eligibility) clearance and eligibility to obtain a CI poly

 

Do you want to join a team that is building tailored technical solutions to modernize our government’s mission and our client’s business?  Do you have a desire to change how people work?  Are you interested in helping to protect our nation’s cyber interests? Join our growing team as a HPC Infrastructure & Cluster Engineer, supporting the NGA customer mission.

Responsibilities What you get to do every day:

 

You will manage the administration, health, and performance of the foundational compute environment. You will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Key Responsibilities:

  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
Qualifications

Clearance Required: Active TS clearance (with SCI Eligibility) and eligibility to obtain CI Poly 

 

Education/Experience:

  • Requires Bachelor's degree
  • 5+years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.

Required Skills:
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand)
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM)
  • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance
  • Troubleshooting Focus: Proven ability to diagnose and resolve complex hardware, network, and OS-level issues

What is ideal?

  • Familiarity with parallel file systems and high-throughput storage architecture.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies
  • Intelligence Community Experience preferred
What you can expect from us
  • Recognizing great achievements do not go unnoticed by bcore through service anniversaries, spot awards, and employee referral bonuses
  • You’ll join a growing organization of passionate, top-shelf, IT engineering professionals with extensive experience in actively developing the technology revolution in the Intelligence community
  • The expected salary range within the Washington, DC metropolitan area is: $148,000 - $179,000. Final compensation is unique to each individual and will be determined based on factors such as experience, education, geographic location, and contractual requirements. This is not a guarantee.
  • Benefits include Health/Dental/Vision, 401(k), Paid Time Off, STD/LTD/Life Insurance/Voluntary Life Insurance, Stipends, Referral Bonuses, and more.

Similar Jobs

More Jobs at Bridge Core (BCore)

More Aerospace & Defense Jobs

Find similar HPC Infrastructure & Cluster Engineer jobs: