GPU DC East-West Network SRE Expert (SME)

Bitdeer Technologies Group

$130K — $160K *
Technical Services
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in data center networking, focused on InfiniBand or RoCE fabrics
  • Hands-on deployment and operation of Nvidia/Mellanox InfiniBand switches at scale
  • Strong understanding of InfiniBand subnet management and QoS
  • Experience with RoCEv2 deployment, including PFC and ECN configuration
  • Proficient with UFM or equivalent IB fabric management tools
  • Knowledge of high-speed optics standards and structured cabling best practices
  • Experience diagnosing network issues using tools like ibdiagnet and perfquery

Responsibilities

  • Manage InfiniBand fabric topologies for large GPU clusters
  • Operate RoCEv2 networks for RDMA workloads across major platforms
  • Monitor and manage IB and RoCE performance and diagnostics
  • Tune NCCL communication based on topology and routing needs
  • Oversee firmware lifecycle for InfiniBand switches and HCAs
  • Diagnose faults and network issues effectively
  • Collaborate with support teams for escalations and troubleshooting

Benefits

  • Flexible work environment
  • Opportunity to impact cutting-edge AI technologies
  • Collaborative team culture
  • Access to advanced training and development programs
  • Opportunity for career advancement
Full Job Description
About the Role

You keep the fabric that makes 10K GPUs act like one - and turn IB/RoCE telemetry into the ground truth for our congestion and link-failure predictors.

Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a healthy training job and a $50M training run stalled by a bad optic. In this role you operate the InfiniBand and RoCEv2 fabrics that carry NCCL traffic across NeoCloud's US DCs, and you feed the AIOps substrate with the fabric telemetry it needs to catch link degradation, congestion, and topology drift before they land on the pager.

What you'll own
  • InfiniBand fabrics: fat-tree, rail-optimized, and dragonfly topologies for GPU clusters of 100-10,000 GPUs.
  • RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads.
  • UFM (Unified Fabric Manager) for IB fabric monitoring, diagnostics, and subnet management.
  • IB and RoCE performance monitoring and tuning: adaptive routing, congestion control (DCQCN/ECN), traffic isolation.
  • NCCL communication tuning: topology detection, ring/tree algorithm selection, GDR configuration.
  • Firmware lifecycle across IB switches and HCAs.
  • Fault diagnosis: link flaps, symbol errors, packet drops, routing anomalies, credit stalls.
  • Coordination with Nvidia/Mellanox support for escalations, bugs, and RMA.

Feed the AIOps substrate
  • Wire IB/RoCE telemetry (ibdiagnet, perfquery, ibstat, PortRcvErrors, PortXmitDiscards, adaptive-routing state) into the platform's collection pipeline.
  • Partner with the platform team to define the Link and Straggler predictors: what a "bad optic 30 minutes from failure" looks like in the counters.
  • Convert every incident into a labeled example the fault-prediction engine can learn from - and every routine mitigation into a workflow the remediation actuator can run.

Job Requirement:
  • 5+ years in data center networking, with at least 3 years focused on InfiniBand or RoCE fabrics
  • Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale
  • Strong understanding of IB subnet management, partitioning, and QoS
  • Experience with RoCEv2 deployment including PFC, ECN, DCQCN configuration
  • Proficiency with UFM or equivalent IB fabric management tools
  • Knowledge of 400G/800G optics, cabling standards, and structured cabling best practices
  • Experience diagnosing IB/RoCE network issues using ibdiagnet, perfquery, ibstat, and similar tools
  • Understanding of NCCL and how GPU communication maps to network topology
  • Instinct for telemetry-driven ops - you've either built dashboards/alerts on RDMA counters at scale, or you can articulate the feature set a fabric-health model would need.
  • Runbook-as-code mindset - the diagnostics you run today should become automation next quarter.

Similar Jobs

More Jobs at Bitdeer Technologies Group

More Technical Services Jobs

Find similar GPU DC East-West Network SRE Expert (SME) jobs: