Job DescriptionAs global demand for AI compute capacity grows at an unprecedented pace, we are building the next generation of hyperscale intelligent computing infrastructure. We are looking for a seasoned network architect with global vision and world-class technical expertise to join our team. You will lead the end-to-end architecture design of AI Data Center Networks (AIDC), high-performance Data Center Interconnect (DCI), and global backbone networks serving large-scale GPU clusters and AI training/inference workloads.
Responsibilities- Network Planning & Architecture: Engage deeply with NVIDIA, leading network equipment vendors, ISPs, and colocation partners; translate business requirements and cloud platform development architecture into top-level designs spanning AIDC internal networks, DCI, and the global backbone. Own the planning and implementation of foundational resources including IP addressing, VLAN/VXLAN, and bandwidth, while charting a forward-looking network evolution roadmap.
- Lead end-to-end AIDC internal network architecture, covering GPU cluster interconnects (InfiniBand / RoCE), Spine-Leaf / Clos topologies, front-end business networks, and out-of-band management networks.
- Collaborate with the cloud platform engineering team on Underlay network architecture design and Overlay network integration (SDN controllers / VxLAN / EVPN, etc.).
- Own congestion control tuning and standard-setting for hyperscale cluster networks, including PFC, ECN, and In-band Network Telemetry (INT).
- Produce architecture documentation (HLD / LLD), network topology diagrams, configuration standards, and SOPs.
- Identify risks in existing network architectures, develop improvement roadmaps, and drive execution to completion.
Qualifications- Bachelor's degree or above in Computer Science, Telecommunications, Electronics, or a related field; 8+ years of network architecture design and planning experience at a hyperscale internet company, top-tier cloud provider, or multinational backbone operator.
- Deep understanding of network requirements for AI compute clusters; expert-level proficiency in InfiniBand (IB) and RoCEv2 architecture design.
- Familiarity with high-performance network topologies such as Rail-Optimized architectures; hands-on experience planning and deploying 10,000- to 100,000-GPU-scale cluster networks is a strong plus.
- Solid command of data center routing and overlay technologies including BGP, OSPF, ECMP, VxLAN, and EVPN; in-depth knowledge and practical experience with Segment Routing (SR-MPLS / SRv6).
- Familiarity with optical transport networks, DWDM technology, and global backbone peering/transit strategies.
- Exceptional cross-functional and cross-cultural communication skills with the ability to align diverse stakeholders on large-scale, complex projects; strong resilience under pressure and a technology-forward mindset.