Mirantis

Senior HPC Networking Engineer

Mirantis$125K — $150K *
US-AnywhereRemote in United States
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in network engineering focusing on HPC or data centers.
  • Strong hands-on experience with InfiniBand technologies like Mellanox/NVIDIA.
  • Solid understanding of networking fundamentals including TCP/IP and routing protocols.
  • Proven experience deploying and troubleshooting Fortinet solutions.
  • Experience with network performance analysis and troubleshooting tools.
  • Familiarity with Linux systems and scripting for automation purposes.
  • Strong analytical and problem-solving skills.

Responsibilities

  • Design, deploy, and maintain HPC network infrastructures focused on InfiniBand.
  • Troubleshoot complex network issues for optimal performance and minimal downtime.
  • Manage and optimize InfiniBand components and fabric configurations.
  • Perform performance tuning, monitoring, and capacity planning for HPC systems.
  • Implement and maintain network security using Fortinet solutions.
  • Collaborate with other teams to support HPC workloads and cluster operations.
  • Develop and maintain documentation for network architecture and procedures.

Benefits

  • Opportunity to work on cutting-edge HPC infrastructure.
  • Collaborative and innovative work environment.
Full Job Description
Job Description

Location: US
Employment Type: Full-time

Role Overview:
We are seeking a highly skilled Senior HPC Networking Engineer to design, deploy, manage, and troubleshoot high-performance networking environments. The ideal candidate will have deep expertise in InfiniBand technologies, strong general networking knowledge, and hands-on experience with Fortinet solutions. You will play a critical role in ensuring the performance, reliability, and scalability of HPC infrastructure.

Key Responsibilities:
  • Design, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics.
  • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
  • Manage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations.
  • Perform performance tuning, monitoring, and capacity planning for HPC networking systems.
  • Implement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).
  • Diagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments.
  • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.
  • Develop and maintain documentation for network architecture, configurations, and operational procedures.
  • Participate in on-call rotations and provide escalation support for critical incidents.
  • Lead or contribute to network upgrades, migrations, and new deployments.
  • You will actively troubleshoot and resolve daily customer incident tickets to keep massive GPU training runs moving.

Build, operate, and scale next-generation GPU infrastructure:
  • You will be hands-on on the front lines driving daily triage, incident resolution, and SLA management for the world's most advanced NVIDIA clusters, InfiniBand/RoCE fabrics, and AI workloads.

Build the playbook, then grow into the platform:
  • Designed for engineers energized by standing up new operations from the ground up, this role offers a direct trajectory from operationalizing bare-metal clusters to driving platform and AI capabilities as we scale.


Qualifications

Required:
  • 5+ years of experience in network engineering, with a focus on HPC or data center environments.
  • Strong hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA).
  • Solid understanding of networking fundamentals: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.
  • Proven experience deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).
  • Experience with network performance analysis and troubleshooting tools.
  • Familiarity with Linux systems and scripting for automation (e.g., Bash, Python).
  • Strong analytical and problem-solving skills.

Preferred:
  • Experience with large-scale HPC clusters or AI/ML infrastructure.
  • Knowledge of RDMA, MPI, and low-latency networking concepts.
  • Certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent.
  • Experience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform).

Soft Skills:
  • Strong communication and collaboration skills.
  • Ability to work independently and handle complex technical challenges.
  • Detail-oriented with a proactive approach to problem-solving.

What We Offer:
  • Opportunity to work on cutting-edge HPC infrastructure.
  • Collaborative and innovative work environment.
  • Competitive salary and benefits package.


Additional Information

About Mirantis

Mirantis is a software company that provides cloud computing services and solutions. The company was founded in 2011 and is headquartered in Sunnyvale, California. Mirantis offers a range of cloud computing services, including OpenStack, Kubernetes, and Docker. The company's solutions are used by a variety of industries, including telecommunications, finance, and healthcare. Mirantis has over 1,000 employees and offices in the United States, Russia, Ukraine, and the United Kingdom.
Learn more about Mirantis
Size
1,000 employees
Industry
Founded
2011

Similar Jobs

More Jobs at Mirantis

More Information Technology Jobs

Find similar Senior HPC Networking Engineer jobs: