Oak Ridge National Laboratory

Senior High Performance Computing Engineer - Classified Environment

Oak Ridge National Laboratory$120K — $145K *
Aerospace & Defense
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • BS in computer science, engineering, or related field with 12+ years of relevant experience.
  • 6+ years in HPC engineering, focusing on cluster management and performance optimization.
  • Experience in classified environments and understanding of security compliance frameworks.
  • Knowledge of HPC systems architecture and cluster management tools like SLURM or PBS.
  • Linux system administration expertise with scripting skills in Bash, Python, or Ansible.
  • Familiarity with benchmarking tools tailored for HPC environments.
  • Experience with parallel programming frameworks such as MPI, OpenMP, or CUDA.

Responsibilities

  • Lead the design and deployment of HPC systems tailored for classified environments.
  • Create and maintain comprehensive documentation of HPC architectures and procedures.
  • Oversee the management and optimization of HPC clusters for performance and reliability.
  • Ensure compliance with security policies through regular audits and implementation of controls.
  • Monitor HPC system performance, identifying and resolving inefficiencies and downtimes.
  • Lead HPC project initiatives and support scientists’ computational needs.
  • Research and drive innovations in HPC technology and infrastructure.

Benefits

  • Collaborative work environment aiming at high security standards.
  • Opportunity to mentor and support the development of junior engineers.
  • Focus on continuous learning and improvement in HPC technologies.
  • Commitment to promoting a respectful workplace culture.
  • Participation in a vibrant team dedicated to customer service and operational excellence.
Full Job Description
Requisition Id 17100

Overview:

Active DOE Q, active DOD Top Secret, or active DOD TS/SCI clearance is required for consideration.

The Field Intelligence Operations Division invites candidates to apply to join the team as a Senior High Performance Computing (HPC) Engineer for Classified Computing to lead the design, implementation, and management of HPC systems within a classified environment. We are looking for candidates with extensive experience in HPC architecture, cluster management, and parallel computing, with a proven ability to work within highly secure and regulated environments. This role involves close collaboration with security teams, scientists, and IT leadership to ensure that the HPC infrastructure meets the stringent performance, security, and compliance requirements necessary for classified work.

As part of our team, you will be joining a vibrant group of professionals eager to provide premier customer service to ensure people and information technology remain secure. The team is collaborative and strives to ensure security practices and procedures are understood, implemented, and enforced.

Major Duties/Responsibilities:
  • HPC System Design and Architecture:
    • Lead the design and deployment of HPC systems, ensuring they meet the computational needs and security requirements of a classified environment.
    • Create and maintain detailed documentation of HPC architectures, configurations, and operational procedures.
  • Cluster Management and Optimization:
    • Oversee the installation, configuration, and management of HPC clusters, ensuring optimal performance, scalability, and reliability.
    • Implement and manage job scheduling, resource allocation, and load balancing to maximize the efficiency of HPC resources.
  • Security and Compliance:
    • Ensure all HPC systems comply with security policies and regulatory requirements, implementing necessary controls and conducting regular audits.
    • Collaborate with the security team to address vulnerabilities and ensure the protection of sensitive data within the HPC environment.
  • Performance Tuning and Troubleshooting:
    • Monitor and optimize the performance of HPC systems, identifying and resolving bottlenecks and inefficiencies.
    • Identify and resolve complex issues, ensuring minimal downtime and disruption to critical operations.
  • Collaboration and Leadership:
    • Lead HPC-related projects, from initial planning and design through to implementation and operational support.
    • Collaborate with scientists, researchers, and others to ensure that the HPC environment meets their computational needs.
    • Mentor and support junior HPC engineers, sharing expertise and best practices.
  • Deliver ORNL's mission by aligning behaviors, priorities, and interactions with our core values of Impact, Integrity, Teamwork, Safety, and Service. Promote equal opportunity by fostering a respectful workplace - in how we treat one another, work together, and measure success.
  • Continuous Improvement and Innovation:
    • Research and remain informed of the latest advancements in HPC technologies, identifying opportunities for innovation and enhancement of the HPC infrastructure.
    • Propose and implement improvements to existing systems and processes to support the evolving needs of the organization.


Basic Qualifications:
  • BS in computer science, engineering, or a related field anda minimum of 12 years of relevant experience. An equivalent combination of education and experience may be considered.
  • 6 years of experience in HPC engineering, with a focus on cluster management, parallel computing, and performance optimization.
  • Demonstrated experience working in classified environments, including a thorough understanding of security policies, compliance frameworks, and associated standard processes (e.g., NIST, DISA STIGs).
  • HPC systems architecture experience, including cluster management tools (e.g., SLURM, PBS, Moab).
  • Linux system administration skills, with experience in scripting and automation using tools such as Bash, Python, or Ansible.
  • Experience with performance tuning and benchmarking tools for HPC environments (e.g., Ganglia, Grafana, or similar).
  • Experience with parallel programming frameworks (e.g., MPI, OpenMP, CUDA) and high-performance interconnects (e.g., InfiniBand).
  • Active DOE Q, active DOD Top Secret, or active DOD TS/SCI clearance is required for consideration.


Preferred Qualifications:
  • Familiarity with advanced storage solutions and parallel file systems (e.g., Lustre, GPFS, or BeeGFS).
  • Professional certifications (e.g., Certified HPC Professional, Linux+, or Security+)
  • Excellent leadership and project management abilities.
  • Strong problem-solving skills with a proactive approach to identifying and resolving issues.
  • Effective communication and collaboration skills, with the ability to work closely with cross-collaborative teams.
  • Ability to manage multiple priorities and work effectively in a fast-paced, high-security environment.
  • Proactive mentality with a commitment to continuous learning and improvement in the rapidly evolving HPC field.


Special Requirements:
  • Visa sponsorship is not available for this position.
  • Work may involve various physical requirements and working conditions.
  • Q clearance with SCI: This position requires the ability to obtain and maintain a Secret Compartmented Information (SCI) clearance from the Department of Energy. As such, this position is a Workplace Substance Abuse (WSAP) testing designated position. WSAP positions require passing a pre-placement drug test and participation in an ongoing random drug testing program. In addition, due the SCI, you may also be subject to random polygraph testing.


Basic: #LI-DNP

This position will remain open for a minimum of 5 days after which it will close when a qualified candidate is identified and/or hired.

We accept Word (.doc, .docx), Adobe (unsecured .pdf), Rich Text Format (.rtf), and HTML (.htm, .html) up to 5MB in size. Resumes from third party vendors will not be accepted; these resumes will be deleted and the candidates submitted will not be considered for employment.

About Oak Ridge National Laboratory

Oak Ridge National Laboratory (ORNL) is a science and technology national laboratory managed for the United States Department of Energy (DOE) by UT-Battelle. ORNL is the largest science and energy national laboratory in the Department of Energy system by size and by annual budget. ORNL conducts research and development activities in a variety of scientific and technical disciplines. ORNL's scientific programs focus on materials, neutron science, energy, high-performance computing, systems biology and national security. ORNL partners with other national laboratories, universities and industry to solve complex problems and transfer knowledge and technology. ORNL is home to several of the world's most powerful supercomputers, including Summit, the world's most powerful supercomputer as of November 2018.
Learn more about Oak Ridge National Laboratory
Size
5,000 employees
Industry
Founded
1943

Similar Jobs

More Jobs at Oak Ridge National Laboratory

More Aerospace & Defense Jobs

Find similar Senior High Performance Computing Engineer - Classified Environment jobs: