HPC Systems Engineer

IREN

$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years experience in HPC system architecture and management.
  • Expertise in Kubernetes integration in HPC environments.
  • Hands-on experience with Slurm workload manager.
  • Familiarity with HPC management tools for monitoring and troubleshooting.
  • Bachelor's or Master's in Computer Science or Engineering.
  • Relevant certifications in Kubernetes or HPC technologies advantageous.
  • Understanding of cloud platforms and HPC integration.

Responsibilities

  • Lead deployment and maintenance of HPC clusters for optimal performance.
  • Integrate and manage components like Kubernetes and Slurm in the HPC setup.
  • Stay updated on HPC and Kubernetes advancements to innovate operations.
  • Collaborate with teams to set best practices for system maintenance.
  • Draft documentation on system designs and operational procedures.
  • Select and integrate management tools for monitoring HPC operations.
  • Provide technical leadership and training to foster team growth.

Benefits

  • 100% company-paid health insurance premiums for employees; 75% for dependents.
  • Company-paid disability insurance and optional life coverage.
  • 401(k) retirement plan with company match.
  • Paid Time Off (PTO) and paid holidays.
  • Internal skills training and professional development opportunities.
  • Access to financial planning and legal services.
  • Company events and team-building activities.
Full Job Description
Job description

Job Type: Full-time | Location: Dallas Fort-Worth, Texas | Department: Information Technologies (IT) | Reporting to: Manager, Systems Engineer | Work Location: #onsite #hybrid

The HPC Systems Engineer will spearhead the design, deployment, and optimization of our high-performance computing (HPC) systems with a focus on HPC clustering, Kubernetes, Slurm, and management tools. The ideal candidate will seamlessly blend traditional HPC solutions with modern container orchestration, ensuring a versatile and scalable computing environment that can be offered to our customers.

Job requirements

  • Minimum of 5 years of experience in HPC system architecture with proven expertise in designing, deploying, and managing HPC clusters.
  • Extensive knowledge of Kubernetes, with a focus on its integration within HPC environments.
  • Hands-on experience with the Slurm workload manager and its intricacies.
  • Familiarity with HPC management tools and software, ensuring efficient system monitoring and troubleshooting.
  • Proven track record of resolving complex system challenges and enhancing operational performance.
  • A Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
  • Relevant certifications in Kubernetes, HPC technologies, or system architecture are advantageous.
  • Understanding of cloud platforms and their integration into HPC ecosystems.
  • Deep knowledge of network and storage solutions commonly used in HPC setups.

Key Attributes:
  • Analytical mindset, adept at envisioning and designing intricate systems.
  • Collaborative approach, with the ability to work effectively with diverse technical teams.
  • Excellent communication skills, translating complex technical concepts into understandable terms for varied audiences.
  • Detail-oriented focus, ensuring systems are both robust and efficient.
  • Continuous learner, keen to stay updated with rapid technological evolutions in the HPC and Kubernetes domains.

Other important requirements:
  • Pre-employment screening, including background check and substance testing may be required according to company policies.
  • Must provide own steel-toed work boots; other PPE will be supplied.
  • Must be able to reliably commute to the work site daily or have plans to relocate before starting work.


Job responsibilities

  • Lead the deployment and maintenance of HPC clusters, ensuring they operate effectively and maximise availability
  • Integrate and manage HPC software components such as Kubernetes, Slurm, cluster management software, and any infrastructure required to operate the HPC environment
  • Stay abreast of advancements in HPC, Kubernetes, and associated technologies, bringing innovations into our operations and product options.
  • Collaborate with technical teams to establish and implement best practices for system maintenance and optimization.
  • Draft comprehensive documentation, including system designs, operational procedures, and best practice guidelines.
  • Facilitate the selection and integration of relevant management tools to monitor, troubleshoot, and enhance HPC operations.
  • Provide technical leadership and training to other team members, fostering an environment of continuous learning and improvement.


Job benefits

The IREN Package

At IREN, we offer a highly competitive compensation package that includes base salary, annual performance incentives, and opportunities to build long-term wealth through equity programs. These offerings are part of our broader Total Rewards package , thoughtfully designed to support your health, well-being, and long-term success.

Compensation
  • Actual compensation will be determined based on factors such as experience, qualifications, and market data for the region.
  • Total Compensation package may be inclusive of annual incentive bonus, equity (long-term incentive).

Health & Wellness
  • 100% company paid health insurance premiums(medical, dental, and vision)for employees, 75% company paid coverage for dependents.
  • Company-paid short-term and long-term disability insurance.
  • Voluntary life, critical illness, and accident coverage available.


  • Health Savings Accounts (HSA) - when combined with the High Deductible Health Plan.
  • Employee Assistance Program and wellness resources.

Retirement & Financial Wealth
  • 401(k) retirement plan with company match.
  • Access to financial planning and legal services.

Time Off & Leave Programs
  • Paid Time Off (PTO) and paid holidays.

Growth & Development
  • Internal skills training and advancement pathways.
  • Professional development to support certifications, continuing education, or role related training.

Community & Culture
  • Company events and team-building activities.

We value diverse perspectives and believe that skills can be developed. If you're passionate about this role, we want to hear from you - whether you meet every criteria or not. Your unique experiences might be exactly what we need!

Similar Jobs

More Information Technology Jobs

Find similar HPC Systems Engineer jobs: