IOC Systems Specialist

IREN

$80K — $95K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2-5 years supporting HPC clusters in production or NOC environments
  • Working knowledge of Kubernetes for workload triage and cluster state interpretation
  • Operational experience with Slurm workload manager for job triage and health checks
  • Familiarity with HPC monitoring tools for proactive issue detection
  • Strong incident response and root cause analysis skills
  • Education in Computer Science or related field; equivalent experience accepted
  • Relevant technical certifications are a plus, such as CKA/CKAD or ITIL

Responsibilities

  • Provide Tier 2 operational support for HPC cloud environments
  • Monitor and troubleshoot incidents across HPC infrastructure
  • Act as escalation point for complex issues, performing root cause analysis
  • Maintain operational monitoring and alerting tools
  • Execute operational changes and maintenance activities
  • Improve operational documentation for efficient IOC operations
  • Identify opportunities for service improvement and automation
  • Support integration of new HPC capabilities into production
  • Mentor Tier 1 staff to enhance IOC knowledge and capability
  • Participate in on-call rotations and major incident responses

Benefits

  • 100% company paid health insurance premiums for employees and 75% for dependents
  • Company-paid short-term and long-term disability insurance
  • Health Savings Accounts (HSA) available
  • 401(k) retirement plan with company match
  • Paid Time Off (PTO) and paid holidays
  • Internal skills training and professional development opportunities
  • Access to financial planning and legal services
  • Company events and team-building activities
Full Job Description
Job description

Job Type: Full-time | Location: Dallas/Fort Worth, TX | Department: Operations | Reporting to: Technical Operations Manager | Work Location Type: #onsite

Job requirements

  • 2-5 years of experience operating or supporting HPC clusters in a production or IOC/NOC environment, including incident triage, change execution, and coordination with engineering teams.
  • Working knowledge of Kubernetes (bare-metal a plus) sufficient to triage workloads, interpret cluster state, and execute documented operational procedures.
  • Operational experience with the Slurm workload manager - job triage, queue and node health checks, and escalation of scheduler issues to engineering.
  • Familiarity with HPC monitoring and observability tooling for alerting, log triage, runbook execution, and proactive issue detection.
  • Demonstrated track record of incident response, root cause analysis, and contributing to operational improvements in complex production systems.
  • Post-secondary education in Computer Science, Engineering, or a related technical field is an asset; equivalent hands-on operational experience is equally valued.
  • Relevant certifications are advantageous - e.g., CKA/CKAD, Linux+/RHCSA, ITIL Foundation, CompTIA Server+, or HPC/GPU vendor certifications.
  • Working understanding of cloud platforms (AWS, Azure, or GCP) and how they integrate with on-premises HPC environments.
  • Working knowledge of the network and storage components common to HPC - InfiniBand/Ethernet fabrics, scale-out storage (e.g., Weka, VAST), and high-throughput interconnects - sufficient to triage issues and engage the right escalation path.


Job responsibilities

The IOC Systems Specialist will focus on the following as well as other job-related duties and/or projects as assigned.
  • Provide Tier 2 operational support for HPC cloud environments, ensuring high availability, performance stability, and adherence to SLAs.
  • Monitor, troubleshoot, and resolve complex incidents across HPC infrastructure, including Kubernetes, Slurm, cluster management systems, and associated cloud services.
  • Act as the escalation point from Tier 1, performing root cause analysis (RCA) and coordinating with Tier 3/engineering teams for defect resolution and permanent fixes.
  • Operate and maintain monitoring, alerting, and observability tooling to proactively detect issues and minimise service disruption.
  • Execute operational changes, patches, upgrades, and maintenance activities in line with change management processes.
  • Maintain and improve operational documentation, including runbooks, playbooks, incident reports, and knowledge base articles to support efficient IOC operations.
  • Contribute to continuous service improvement by identifying recurring issues, operational risks, and automation opportunities within the HPC environment.
  • Support tooling integration and operational readiness for new HPC capabilities prior to handover into production support.
  • Provide technical guidance and mentoring to Tier 1, DC Techs, enhancing IOC capability and knowledge depth.
  • Participate in on-call rotations and major incident response activities.


Job benefits

At IREN, we offer a highly competitive compensation package that includes base salary, annual performance incentives, and opportunities to build long-term wealth through equity programs. These offerings are part of our broader total rewards package, thoughtfully designed to support your health, well-being, and long-term success.

Compensation
  • Actual compensation will be determined based on factors such as experience, qualifications, and market data for the region.
  • Total Compensation package may be inclusive of annual incentive bonus, equity (long-term incentive).
Health & Wellness
  • 100% company paid health insurance premiums (medical, dental, and vision) for employees, 75% company paid coverage for dependents
  • Company-paid short-term and long-term disability insurance
  • Voluntary life, critical illness, and accident coverage available
  • Health Savings Accounts (HSA) - when combined with the High Deductible Health Plan
  • Employee Assistance Program and wellness resources
  • Retirement & Financial Wealth
  • 401(k) retirement plan with company match
  • Access to financial planning and legal services
Time Off & Leave Programs
  • Paid Time Off (PTO) and paid holidays
Growth & Development
  • Internal skills training and advancement pathways
  • Professional development to support certifications, continuing education, or role related training
Community & Culture
  • Company events and team-building activities

We value diverse perspectives and believe that skills can be developed. If you're passionate about this role, we want to hear from you - whether you meet every criteria or not. Your unique experiences might be exactly what we need!

Similar Jobs

More Information Technology Jobs

Find similar IOC Systems Specialist jobs: