Senior Systems Engineer - HPC & GPU Infrastructure

Base-2 Solutions, LLC

• $120K — $145K *
Technical Services
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • 12+ years of systems engineering experience in computer science or electrical engineering fields.
  • Expertise in Linux operating system integration.
  • Strong knowledge of computer hardware architecture in Linux systems.
  • Proficiency in scripting languages like Python or BASH.
  • Experience with automation tools such as Ansible or Terraform.
  • Valid DoD 8570.11 IAT Level II certification required.

Responsibilities

  • Install and maintain GPU and HPC hardware on-premises and cloud environments.
  • Analyze and improve performance of HPC and GPU clusters in Linux.
  • Configure job-scheduling platforms like Slurm and Kubernetes.
  • Apply power-management techniques for GPU efficiency.
  • Design and execute Linux-based tests for GPU functionality and performance.
  • Document architectural specifications and best practices for GPU development.
  • Research and propose Linux-based innovations in GPU optimization.

Benefits

  • 100% company-paid medical, dental, and vision premiums for employees and dependents.
  • Competitive 401(k) with 4% company match and immediate vesting.
  • Up to 20 days of flexible paid time off (PTO) plus 11 paid holidays.
  • Flexible work schedules for improved work-life balance.
  • Employee referral bonuses up to $10,000.
Full Job Description
  • Requisition ID: 5065
  • Standard Title:
  • Required Security Clearance: Top Secret/SCI
  • Location: Bethesda, MD
  • Work Type: On-Site
  • Shift: First
  • Referral Eligibility: Eligible
  • U.S. Citizenship Required? Yes


Position Summary

Base-2 Solutions is seeking a Senior Systems Engineer - HPC & GPU Infrastructure to design, develop, and optimize high-performance computing and GPU clusters supporting Intelligence Community customers. This is a 100% on-site position at the customer site in Bethesda, Maryland.
Essential Duties and Responsibilities
  • Contribute to the installation and maintenance of GPU and HPC hardware on-premises and in the cloud, providing insight into hardware performance and interaction with software components.
  • Analyze HPC and GPU cluster performance, identify bottlenecks, and develop strategies to improve performance across applications in Linux, addressing hardware and software considerations.
  • Install and configure HPC and GPU job-scheduling and workload-management platforms, including Slurm, PBS, Apache Airflow, and Kubernetes.
  • Apply power-management techniques to optimize GPU power consumption on mobile and desktop Linux platforms and continuously assess power-efficiency strategies.
  • Design and execute Linux-based tests to validate GPU performance and functionality, including stress testing, benchmarking, and debugging; maintain and expand the testing suite.
  • Maintain comprehensive technical documentation, including architectural specifications, code documentation, and Linux-specific best practices for GPU development.
  • Monitor trends, innovations, and competitive developments in the GPU industry; contribute to research and propose Linux-specific approaches to GPU design and optimization.
Required Qualifications
  • Relevant systems engineering experience meeting an applicable education and experience pathway in the equivalency section.
  • Expertise in operating system integration for Linux.
  • Strong understanding of computer hardware architecture, particularly as it relates to Linux systems.
  • Knowledge of parallel computing, graphics algorithms, and real-time rendering in Linux environments.
  • Excellent problem-solving skills and ability to collaborate within a team.
  • Strong communication skills for conveying technical information in a Linux context.
  • Proficiency with scripting languages such as Python or BASH.
  • Proficiency with automation tools such as Ansible, Puppet, Salt, Terraform, and similar tools.
  • Candidate must, at a minimum, meet DoD 8570.11 IAT Level II certification requirements; an IAT Level III certification is also acceptable.
  • US Citizenship is required due to the nature of the government contracts supported.
Preferred Qualifications
  • Knowledge of GPU virtualization, cloud computing, and emerging Linux-based technologies.
  • Experience with container technologies, including Docker and Kubernetes.
  • Experience with Prometheus/Grafana for monitoring.
  • Knowledge of distributed resource scheduling systems.
  • Understanding of data center networking hardware and cabling concepts.
  • Understanding of networking technologies such as DHCP, DNS, TCP/IP, VLANs, HSRP, and SNMP.
  • Knowledge of data center networking security principles, including firewall ACLs, IPS/IDS, and policy-based routing.
Required Education and Experience Equivalency
  • Bachelor's or higher degree in Computer Science, Electrical Engineering, or a related field with 12+ years of relevant systems engineering experience.
  • Additional years of relevant systems engineering experience in lieu of a degree.
Required Certifications
  • Candidate must, at a minimum, meet DoD 8570.11 IAT Level II certification requirements, currently Security+ CE, CCNA-Security, GICSP, GSEC, or SSCP along with an appropriate computing environment (CE) certification. An IAT Level III certification would also be acceptable, including CASP+, CCNP Security, CISA, CISSP, GCED, GCIH, or CCSP.
Required Security Clearance
  • Active Top Secret/SCI


Pay & Benefit Highlights
Compensation
  • Competitive fixed salary or hourly pay (based on experience, skills, location, and internal equity).
  • Employee referral bonuses up to $10,000 per hired referral.
  • Additional bonus opportunities for exceptional performance and contributions to business development and company growth (role-dependent).
Health
  • 100% company-paid medical premiums for employees and eligible dependents.
  • Choose from multiple plan options with CareFirst, Kaiser, and UnitedHealthcare, including PPO, POS, HMO, and HSA-compatible plans.
  • 100% company-paid dental premiums for employees and eligible dependents.
  • 100% company-paid vision premiums for employees and eligible dependents.
Income Protection
  • 100% company-paid premiums for short-term disability.
  • 100% company-paid premiums for long-term disability.
  • 100% company-paid premiums for accidental death & dismemberment (AD&D).
  • 100% company-paid premiums for life insurance up to $200,000.
Retirement
  • 401(k) with immediate vesting: 4% company match plus a 4% non-elective company contribution (8% total).
  • 401(k) pre-tax and Roth options.
Leave
  • Up to 20 days of flexible paid time off (PTO).
  • 11 paid floating holidays.
Work-Life Balance
  • Flexible work schedules, including flex time and compressed work periods (contract and project-dependent).


NT

Similar Jobs

More Jobs at Base-2 Solutions, LLC

More Technical Services Jobs

Find similar Senior Systems Engineer - HPC & GPU Infrastructure jobs: