DescriptionSystems Engineer - HPC and GPU Infrastructure - TS/SCI As a Systems Engineer - HPC & GPU Infrastructure, you will play a pivotal role in designing, developing, and optimizing GPU clusters for the IC community customers.
Location:Bethesda, MD
Security Clearance:TS/SCI and willingness to get a Polygraph or TS/SCI with CI Polygraph
Responsibilities:- HPC and GPU environment engineering: Contribute to the installation and maintenance of GPU and HPC hardware on-prem and in the cloud, providing insights into hardware performance to ensure efficient interaction with software components.
- Performance Optimization: Analyze HPC/GPU cluster performance, identify bottlenecks, and develop strategies to enhance performance across various applications in Linux, addressing both hardware and software considerations. Regularly monitor and improve performance.
- HPC/GPU tooling: Install and configure HPC/GPU job scheduling and workload management platforms such as Slurm, PBS , Apache Airflow, Kubernetes
- Power Efficiency: Work on power management techniques to optimize GPU power consumption, ensuring efficient operation on both mobile and desktop Linux platforms. Continuously assess and enhance power efficiency strategies.
- Testing and Validation: Design and execute tests to validate GPU performance and functionality on Linux, including stress testing, benchmarking, and debugging to ensure robust operation. Maintain and expand the testing suite.
- Documentation: Maintain comprehensive technical documentation, including architectural specifications, code documentation, and Linux-specific best practices for GPU development. Keep documentation up to date with changes and improvements.
- Industry Insight: Stay updated on the latest trends, innovations, and competitive landscapes within the GPU industry, contributing to research efforts and proposing Linux-specific approaches to GPU design and optimization. Share regular updates and insights with the team
Minimum Requirement: - Bachelor's or higher degree in Computer Science, Electrical Engineering, or a related field. Additional years of experience may be considered in lieu of a degree.
- 4+ years of relevant systems engineering experience
- Expertise in operating system integration for Linux.
- Strong understanding of computer hardware architecture, particularly as it relates to Linux systems.
- Knowledge of parallel computing, graphics algorithms, and real-time rendering in Linux environments.
- Excellent problem-solving skills and the ability to collaborate within a team.
- Strong communication skills for conveying technical information in a Linux context.
- Proficiency with scripting languages such as Python or BASH.
- Proficiency with automation tools such Ansible, Puppet, Salt, Terraform, etc.
- Candidate must, at a minimum, meet DoD 8570.11- IAT Level II certification requirements (currently Security+ CE, CCNA-Security, GICSP, GSEC, or SSCP along with an appropriate computing environment (CE) certification). An IAT Level III certification would also be acceptable (CASP+, CCNP Security, CISA, CISSP, GCED, GCIH, CCSP).
Preferred Qualifications: - Knowledge of GPU virtualization, cloud computing, and emerging Linux-based technologies in the field.
- Experience with container technologies (Docker, Kubernetes)
- Experience with Prometheus/Grafana for monitoring
- Knowledge of distributed resource scheduling systems
- Understanding data center networking hardware and cabling concepts.
- Understanding of networking technologies such as DHCP, DNS, TCP/IP, VLANs, HSRP, and SNMP.
- Knowledge of data center networking security principles Firewall ACLs, IPS/IDS, and Policy Based Routing.