Senior AI Data Center Network Engineer

Bitdeer Technologies Group

$130K — $155K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Network Engineering, Telecommunications, or related field
  • 8-10 years of experience in data center network operations, architecture, and engineering
  • Expertise in TCP/IP protocol stack and routing protocols (BGP, OSPF, ISIS)
  • Hands-on experience with HPC/AI networking: InfiniBand architecture and RoCEv2
  • Proficiency in managing data center switches and routers from vendors like Cisco and NVIDIA/Mellanox
  • Strong skills in network automation with Python, Ansible, Terraform; Linux admin experience
  • Fluency in English and Chinese for global collaboration

Responsibilities

  • Architect scalable network solutions for AI Cloud Data Centers using advanced topologies and technologies
  • Design and optimize large-scale GPU clusters with InfiniBand and RoCEv2 fabrics
  • Lead troubleshooting for network issues impacting AI workloads
  • Develop and maintain network automation platforms for streamlined configuration management
  • Utilize monitoring tools to ensure network health and performance
  • Support Kubernetes/Docker networking in cloud environments
  • Lead network changes and incident response efforts, conducting Root Cause Analysis

Benefits

  • Collaborate with cutting-edge technology in AI and cloud infrastructure
  • Opportunity to work with a global team and impact major AI workloads
  • Engagement with advanced networking technologies
  • Emphasis on continuous learning and professional growth
  • Involvement in high-performance, high-availability projects
Full Job Description
Position Overview
We are seeking an experienced and highly skilled Senior AI Data Center Network Engineer to join the Bitdeer AI Cloud team in the US. In this role, you will architect, deploy, and operate high-performance, high-availability network solutions for our large-scale GPU clusters. You will be instrumental in ensuring the stability, low latency, and optimal performance of our AI infrastructure, focusing heavily on advanced networking technologies such as InfiniBand, RoCEv2, and SDN architectures. You will collaborate with cross-functional teams to build the backbone of our AI Cloud, enabling cutting-edge AI training and inference workloads for our global customers.

Key Responsibilities
  • Network Architecture & Design: Architect scalable, high-availability network solutions for AI Cloud Data Centers, including Spine-Leaf topology, DCN, DCI, and backbone networks using VXLAN EVPN and SDN technologies.
  • High-Performance AI Networking: Design, deploy, and optimize large-scale GPU clusters utilizing InfiniBand and RoCEv2 (Lossless Ethernet) fabrics. Configure and manage NVIDIA Spectrum/Quantum series switches.
  • Cluster Operations & Troubleshooting: Lead deep-dive investigations into complex network issues affecting AI workloads, such as RDMA packet loss, congestion control (PFC/ECN), latency, and NCCL communication timeouts.
  • Network Automation & DevOps: Develop and maintain network automation tools and platforms using Python, Ansible, and Terraform to implement Infrastructure as Code (IaC) and streamline configuration management.
  • Fabric Management & Monitoring: Utilize tools like NVIDIA UFM (Unified Fabric Manager), NetQ, Zabbix, and Prometheus to ensure real-time monitoring, network telemetry, and overall fabric health.
  • Cloud & Container Networking: Support network integration for Kubernetes/Docker container environments and hybrid cloud deployments, ensuring seamless connectivity and security.
  • Incident Response & Change Management: Lead critical network changes, capacity expansions, and firmware upgrades. Provide rapid response to incidents, implement mitigations, and conduct Root Cause Analysis (RCA).

Requirements
  • Bachelor's degree or above in Computer Science, Network Engineering, Telecommunications, or a related field.
  • Minimum 8-10 years of experience in large-scale data center network operations, architecture, and engineering.
  • Deep expertise in the TCP/IP protocol stack and core routing protocols (BGP, OSPF, ISIS), as well as Data Center technologies (VXLAN EVPN, Spine-Leaf).
  • Extensive hands-on experience with HPC/AI networking: InfiniBand architecture, Subnet Manager, or Ethernet-based RoCEv2, including PFC, ECN, and congestion control mechanisms.
  • Proficiency in configuring and managing data center switches and routers from major vendors (e.g., Cisco Nexus, NVIDIA/Mellanox, Juniper).
  • Strong skills in network automation and scripting (Python, Ansible, Terraform) and Linux system administration.
  • Experience with Kubernetes container networking (CNI) and cloud-native architectures.
  • Professional fluency in English and Chinese is required to effectively collaborate with global teams.

Similar Jobs

More Jobs at Bitdeer Technologies Group

More Information Technology Jobs

Find similar Senior AI Data Center Network Engineer jobs: