Job Summary
We are seeking a senior GPU Network Engineer to provide Tier 3 operational support for GPU networking environments supporting AI and machine learning workloads. This role will focus on troubleshooting complex issues involving GPU clusters, high-performance networking, RoCE, InfiniBand, and AI infrastructure. The engineer will serve as a technical escalation point and collaborate with network, compute, platform, security, and operations teams to ensure performance, availability, and rapid resolution of production issues. The ideal candidate will have hands-on experience supporting GPU clusters, AI/ML infrastructure, High Performance Computing (HPC) environments, or large-scale low-latency data center networks.
Key Responsibilities
• Serve as a Tier 3 escalation resource for GPU network incidents and production issues.
• Troubleshoot complex networking, performance, and connectivity issues impacting GPU clusters.
• Support AI, machine learning, and high-performance computing environments.
• Analyze network performance, congestion, latency, packet loss, and throughput issues.
• Collaborate with security operations, network, compute, platform, and infrastructure teams during incident response.
• Validate and optimize GPU network performance for customer workloads.
• Identify root causes and implement corrective actions for recurring issues.
• Support production turn-up and go-live activities for new GPU environments.
• Participate in operational readiness reviews and knowledge transfer sessions.
• Develop operational runbooks, troubleshooting guides, and support procedures.
• Provide recommendations for capacity planning, scalability, and operational improvements.
Required Qualifications
• 5+ years of data center networking experience.
• 2+ years of experience supporting GPU, AI/ML, HPC, or large-scale compute infrastructure environments.
• Experience troubleshooting complex network performance issues in production environments.
• Strong understanding of data center network architecture and operations.
• Experience supporting high-bandwidth, low-latency network fabrics.
• Experience with Arista, NVIDIA, Cisco, or equivalent data center networking technologies.
• Strong understanding of Layer 2 and Layer 3 networking.
• Strong understanding of routing and switching.
• Experience with network monitoring and troubleshooting.
• Experience with incident management and escalation processes.
• Ability to work effectively during critical outages and customer-impacting incidents.
Preferred Qualifications
• Experience supporting NVIDIA DGX, HGX, SuperPOD, or similar GPU infrastructure.
• Experience with RDMA, RoCE, InfiniBand, AI network fabrics, or east-west traffic optimization.
• Knowledge of NVIDIA reference architectures.
• Experience with Kubernetes-based AI environments.
• Experience supporting hyperscale, cloud, NeoCloud, or HPC environments.
• Exposure to optical networking, buffer tuning, congestion management, and performance validation.
• Familiarity with network telemetry and performance analytics tools.
_:empty]:hidden">