Position OverviewWe are reinforcing the network team that underpins a multi-tenant, multi-region GPU cloud at 10,000+ GPU scale. You will design and operate tenant network isolation and high-performance GPU interconnect, ensuring low-latency, lossless fabrics for distributed AI training alongside secure multi-tenancy.
Key Responsibilities- Design, operate, and troubleshoot multi-tenant SDN: tenant isolation, VPC / overlay networking, and east-west / north-south traffic.
- Build and tune high-performance GPU interconnect fabrics (InfiniBand / RoCE) for distributed training, including lossless config and congestion control (PFC / ECN).
- Plan and operate cross-region interconnect and routing.
- Lead network change management and fast fault localization to minimize tenant impact.
- Partner with Compute and Control Plane teams on network-provisioning automation.
- Define network architecture standards, security boundaries, and observability.
Job Requirement:- 6+ years in datacenter / cloud networking or SDN.
- Deep expertise in multi-tenant isolation, overlays (VXLAN / EVPN), and L2 / L3 design.
- Hands-on with high-performance fabrics: InfiniBand and / or RoCE, RDMA - strongly preferred for GPU / HPC clusters.
- Experience with SDN controllers and network automation.
- Strong troubleshooting / fault-localization under production load.
- Multi-region network design experience a plus.