Advanced Micro Devices, Inc

Senior Network Engineer - GPU Cluster Networking

Advanced Micro Devices, Inc$145K — $175K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in network engineering, particularly with backend networks for GPU clusters
  • Deep expertise in data center networking concepts and technologies
  • Hands-on experience with RDMA and RoCEv2 in large production environments
  • Strong knowledge of modern network architectures, including leaf-spine and fat-tree
  • Proficient in configuring and troubleshooting network protocols such as BGP, ECMP, and VLANs
  • Bachelor's or Master's degree in Computer Engineering or related field, or equivalent experience

Responsibilities

  • Architect and deploy high-performance backend networks for AMD Instinct GPUs
  • Design scalable network fabrics for AI and HPC environments with 10,000+ GPUs
  • Manage the network path from GPU server through the switching fabric
  • Optimize high-speed Ethernet fabrics using RoCEv2 and multiple GbE technologies
  • Conduct network topology modeling and traffic-flow analysis
  • Lead incident responses and implement corrective actions
  • Plan and execute network expansions and upgrades for GPU clusters

Benefits

  • Comprehensive health and wellness programs
  • Retirement savings plan
  • Flexible work environment, hybrid options available
  • Access to ongoing training and professional development
  • Opportunities to collaborate on cutting-edge technology projects
Full Job Description
THE ROLE:

We are seeking a Senior Network Engineer to join the AMD IT System Engineering team.

This role is responsible for the architecture, deployment, optimization, automation, and production operation of high-performance backend networks supporting large-scale AMD GPU clusters. The engineer will own the network path from the GPU server and NIC through the data center switching fabric, ensuring that distributed AI training, large language model, inference, and HPC workloads receive predictable bandwidth, low latency, and reliable collective communication performance.

The ideal candidate will have experience designing, scaling, and operating backend network infrastructure for GPU clusters with approximately 10,000 or more GPUs, or comparable hyperscale AI and HPC environments.

The primary focus of this position is high-speed Ethernet and RoCEv2 networking for AMD Instinct accelerator clusters. You will work across switches, NICs, optics, RDMA, Linux networking, PCIe and NUMA topology, ROCm, RCCL, SLURM, Kubernetes, storage networks, automation platforms, and observability systems.

You will partner with AMD AI engineering, network engineering, data center, storage, security, platform, and application teams to ensure the backend network fabric is not a bottleneck to GPU workload performance.

THE PERSON:

You are a highly experienced, hands-on network engineer with deep expertise in data center networking, RDMA, RoCEv2, and large-scale GPU cluster fabrics with approximately 10,000 or more GPUs,.

You understand how distributed GPU workloads generate traffic across the backend network and how application performance is affected by network topology, congestion, GPU-to-NIC locality, routing, switch buffering, traffic-class configuration, and collective communication patterns. You take responsibility for end-to-end outcomes, including architecture, implementation, qualification, production deployment, monitoring, incident response, capacity planning, and continuous improvement. You use telemetry and repeatable performance testing to validate designs and make data-driven engineering decisions.

You are comfortable leading complex technical initiatives, mentoring engineers, documenting architecture and operating standards, and working across globally distributed organizations.

KEY RESPONSIBILITIES:
  • Architect, deploy, operate, and continuously improve high-performance backend networks for large-scale AMD Instinct GPU clusters.
  • Design network fabrics capable of supporting AI and HPC environments ranging from individual GPU racks to clusters containing 10,000 or more GPUs.
  • Own the backend network architecture from the GPU server and network interface card through the leaf-spine switching fabric.
  • Design and optimize high-speed Ethernet fabrics using RoCEv2 and 100/200/400 GbE technologies.
  • Develop scalable network topologies, including leaf-spine, Clos, fat-tree, rail-optimized, multi-plane, and non-blocking fabric architectures.
  • Perform network topology modeling, oversubscription analysis, traffic-flow analysis, bandwidth planning, port-capacity planning, failure-domain analysis, and long-term growth forecasting.
    • Configure, tune, validate, and troubleshoot lossless or near-lossless RoCEv2 environments, including PFC, ECN, DCQCN, QoS, ECMP, Switch buffer and queue management, DSCP and priority mappings
  • Design and operate routing and switching environments using technologies such as BGP, ECMP, VLAN, VRF, EVPN, and VXLAN.
  • Optimize end-to-end communication performance across GPUs, NICs, switches, CPUs, PCIe devices, storage systems, and the Linux networking stack.
  • Lead production incident response, root-cause analysis, corrective actions, and preventive engineering improvements for GPU cluster networks.
  • Plan and execute network expansions, cluster scale-outs, switch replacements, capacity upgrades, and fabric migrations

PREFERRED EXPERIENCE:
  • Significant experience designing, deploying, and operating production data center networks for AI, GPU, HPC, cloud, or other large-scale distributed computing environments.
  • Experience designing, scaling, or operating backend network infrastructure for GPU clusters containing approximately 10,000 or more GPUs, or similarly sized hyperscale compute environments.
  • Deep knowledge of data center networking fundamentals; Routing and switching, VLANs and subnetting, BGP and ECMP, Quality of Service, MTU configuration, Switch buffering, Network segmentation
  • Strong hands-on experience with RDMA and RoCEv2 in production environments.
  • Demonstrated experience configuring, tuning, and troubleshooting PFC, ECN, DCQCN, QoS, switch buffers, NIC queues, RDMA traffic classes, and lossless or near-lossless Ethernet.
  • Strong understanding of leaf-spine, Clos, fat-tree, rail-optimized, and multi-plane network architectures.
  • Experience with network routing technologies such as BGP and ECMP and overlay technologies such as EVPN and VXLAN.
  • Strong understanding of GPU cluster topology, including GPU-to-GPU, GPU-to-NIC, CPU-to-NIC, PCIe, NUMA, and network locality.
  • Experience building monitoring and observability solutions using Prometheus, Grafana, streaming telemetry, gNMI, SNMP, sFlow, or equivalent platforms.
  • Experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting
  • Experience with AMD Instinct accelerators, ROCm, RCCL, and AMD GPU software environments.
  • Experience with AMD Pensando AI NICs, SmartNICs, DPUs, or other AMD Pensando networking technologies.
  • Strong hands-on experience with Juniper data center switching platforms and Junos OS, including configuration and troubleshooting
  • Experience designing backend networks specifically for large language model training and other communication-intensive distributed AI workloads.
  • Experience with Ethernet fabric technologies such as BGP, EVPN, VXLAN, and modern leaf-spine data center architectures.

ACADEMIC CREDITALS:
  • Bachelor's or Master's degree in Computer Engineering, or a related field, or equivalent practical experience.

LOCATION:

San Jose, CA

This role is not eligible for visa sponsorship.

#LI-MF2

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

About Advanced Micro Devices, Inc

Advanced Micro Devices, Inc. Careers

Join the innovative forefront of technology with a career at Advanced Micro Devices, Inc. (AMD), a leader in semiconductor development. As part of our global team, you will contribute to an organization renowned for its dedication to innovation, leadership, and diversity in the tech industry.

Work You’ll Do

At AMD, we offer job opportunities that push the boundaries of what is possible. Our team is composed of professionals who lead the way in microprocessor and graphics technology, driving industry standards and innovation. With AMD, you will be part of a culture that values growth and professional development, ensuring that every team member has the opportunity to excel.

Transform Your Career

AMD is not just about advancing technology, but also about advancing careers. Whether you are looking for an internship, a full-time position, or leadership roles, AMD provides the platform to propel your career to new heights. Our commitment to professional growth is matched by our dedication to diversity and inclusion, making AMD a place where everyone can thrive.

Innovative Work Environment

Join a team of over 12,000 dedicated professionals at the intersection of technology, industry expertise, and digital innovation. At AMD, you will work on groundbreaking projects that shape the future of computing and graphics. Our collaborative environment encourages networking and the sharing of ideas across teams and disciplines.

Career Development and Benefits

AMD is committed to the development of its employees. We offer robust training programs, including leadership development and diversity training, to ensure our team is equipped for both current challenges and future opportunities. Our benefits package is designed to support the well-being and financial security of our employees and their families.

Explore Job Opportunities

From engineering to marketing, AMD offers a range of career paths that cater to diverse skills and interests. Our hiring process is designed to be transparent and engaging, helping you to understand where you fit within our team and how you can contribute to our collective goals.

Stay Connected

Join Our Team Search open positions that match your skills and interest. We look for passionate, curious, creative, and solution-driven team players. Explore the opportunities to join a company that’s committed to your career growth and to innovation in the technology sector.

Keep Up to Date

Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here.

Job Alert Emails

Personalize your subscription to receive job alerts, latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities that await at Advanced Micro Devices, Inc.

Interview and Resume Tips

Prepare for your future with AMD by accessing resources that help you craft your resume and excel in interviews. Our goal is to help you showcase your best professional self and align your skills with the needs of our dynamic team. At Advanced Micro Devices, Inc., we empower our employees to innovate, lead, and grow. Join us in driving the future of technology while building a rewarding and sustainable career.
Learn more about Advanced Micro Devices, Inc
Size
15,500 employees
Market Cap
$100.9 billion
Industry
Net Income
$2.4 billion
Founded
1969
5 Year Trend
+30.9%
Revenue
$9.7 billion
NASDAQ

Similar Jobs

More Jobs at Advanced Micro Devices, Inc

More Information Technology Jobs

Find similar Senior Network Engineer - GPU Cluster Networking jobs: