12+ years in large-scale compute, HPC, or GPU infrastructure engineering with production experience
Deep understanding of GPU interconnect and system topology including NVLink and PCIe
Strong RDMA networking background with RoCEv2 or InfiniBand expertise
Practical NCCL expertise in debugging and optimizing performance
Substantial Kubernetes experience for accelerated workloads
Strong programming skills in Go and Python, with knowledge of driver and firmware layers
Experience managing hardware failure diagnostics and automation
Responsibilities
Own the end-to-end reference architecture for rack-scale inference systems
Define NVLink domain partitioning and allocation across tenants
Architect the scale-out RDMA fabric connecting racks
Make Kubernetes topology-aware for GPU scheduling
Drive multi-node disaggregated serving in partnership with the inference platform team
Own failure semantics and serviceability at rack scale
Establish validation and burn-in methodology for production qualification
Collaborate with data center engineering on power and cooling constraints
Work with NVIDIA on pre-silicon enablement and hardware bringup
Set technical direction and mentor engineers across teams
Benefits
Hybrid work model
Opportunity to work with cutting-edge GPU technology
Collaboration with industry leaders like NVIDIA
Mentorship and leadership opportunities
Involvement in shaping the future of inference architecture
Full Job Description
We are looking for a Principal Engineer to own how rack-scale GPU systems are architected, orchestrated, and operated for inference at DigitalOcean.
DigitalOcean is the Inference Cloud. The unit of GPU capacity is changing underneath the entire industry: for a decade the server was the boundary, and now it is the rack. GB300 NVL72 and the Vera Rubin generation behind it present 72 or more GPUs inside a single coherent NVLink domain, which breaks most of the assumptions our scheduling, networking, failure-handling, and capacity models were built on. Getting this right determines whether we can serve trillion-parameter models with the economics our customers expect.
This role owns that transition. You will define the reference architecture for rack-scale inference at DigitalOcean-how NVLink domains are carved and allocated, how the scale-out RDMA fabric is designed and tuned around them, how Kubernetes is taught to schedule against topology it was never designed to understand, and what happens when one GPU in a 72-GPU coherent domain fails at 3am. You'll work across our data center, networking, and inference platform teams, and directly with NVIDIA on pre-silicon enablement and bringup. What You'll Be Doing:
Owning the end-to-end reference architecture for rack-scale inference systems-GB300 NVL72 today, Vera Rubin-class NVL144 and successors next-from fabric design through to what a customer can actually schedule
Defining how NVLink domains are partitioned, allocated, and isolated across tenants and workloads, including multi-node NVLink and IMEX domain provisioning, fabric manager configuration, and the blast radius each choice implies
Architecting the scale-out RDMA fabric that connects racks-RoCEv2 or InfiniBand, rail-optimized topologies, congestion control, GPUDirect RDMA paths-and tuning NCCL to actually use it well
Making Kubernetes topology-aware for these systems: DRA for GPU and NVLink-domain allocation, multi-node inference primitives such as LeaderWorkerSet, gang scheduling, Topology Manager and NUMA alignment, and the GPU and Network Operators that underpin all of it
Driving multi-node disaggregated serving on this hardware in partnership with the inference platform team-how prefill and decode pools map onto NVLink domains and rail boundaries, and where the KV cache moves between them
Owning failure semantics and serviceability at rack scale: health checking and DCGM-based diagnostics, drain and repair workflows, degraded-domain scheduling, and firmware and driver lifecycle across a fleet of racks
Establishing the validation and burn-in methodology that qualifies a rack for production-collective bandwidth and latency characterization, NCCL and end-to-end inference benchmarks, and the acceptance gates a rack must pass
Partnering with our data center engineering team on power, liquid cooling, and rack density constraints, and translating those into what is actually deployable and at what cost
Working directly with NVIDIA and ODM partners on pre-silicon enablement, early hardware bringup, and roadmap feedback
Setting technical direction, mentoring senior and staff engineers across teams, and representing DigitalOcean upstream, at conferences, and in customer architecture reviews
What You'll Add to DigitalOcean:
12+ years in large-scale compute, HPC, or GPU infrastructure engineering, including hands-on responsibility for systems in production
Deep understanding of GPU interconnect and system topology-NVLink and NVSwitch, NVLink domains, PCIe, NUMA, and how each shapes what a workload can be placed where
Strong RDMA networking background: RoCEv2 or InfiniBand at scale, GPUDirect RDMA, congestion control and lossless fabric tuning, and rail-optimized cluster topologies
Practical NCCL expertise-not just running it, but debugging it: algorithm and protocol selection, topology detection, and diagnosing collectives that are slow for non-obvious reasons
Substantial Kubernetes depth for accelerated workloads: device plugins and DRA, scheduling extensions, operators, and a clear-eyed view of where Kubernetes' model breaks down against rack-scale hardware
Strong systems programming and automation skills in Go and Python, and comfort at the driver, firmware, and BMC layer when the problem lives there
Experience operating fleets where hardware failure is routine, including the diagnostics, automation, and blast-radius thinking that makes that survivable
Excellent written and verbal communication, and a track record of leading architecture across organizational boundaries-hardware, networking, platform, and product
Bonus:
Direct experience bringing up GB200 or GB300 NVL72 systems, or comparable rack-scale platforms, in a production environment
Experience with multi-node inference for very large models-expert or pipeline parallelism spanning NVLink domains, and the placement constraints that follow
Familiarity with liquid-cooled and high-density rack deployments, and the operational realities of servicing them
Contributions to relevant open source: Kubernetes SIG-Node or scheduling, NVIDIA GPU or Network Operator, DRA drivers, Kueue, LeaderWorkerSet, or NCCL
Experience running heterogeneous fleets spanning multiple accelerator vendors and generations under a single orchestration model