Position OverviewAs the platform scales toward a 10,000+ GPU, multi-region footprint, we are expanding the team that turns procured GPU capacity into reliable, sellable compute. You will own the end-to-end lifecycle of bare-metal GPU nodes-from provisioning and delivery through break-fix, firmware/DPU management, and decommissioning-directly driving delivery throughput, fleet availability, and the ROI of our largest capital investment.
Key Responsibilities- Own the full lifecycle of bare-metal GPU nodes: provisioning, delivery / onboarding, in-service operation, break-fix, and decommissioning across multiple regions.
- Build and operate automated, repeatable node-delivery pipelines to eliminate the current delivery backlog and keep pace with fleet growth toward 10,000+ GPUs.
- Manage DPU / SmartNIC and server firmware (BMC / BIOS / NIC / GPU firmware): version baselines, upgrades, and validation.
- Drive fleet reliability: reduce MTTR, lead incident response and root-cause analysis, and improve hardware-health monitoring.
- Participate in a sustainable 7x24 multi-region on-call rotation; build runbooks and tooling that reduce manual toil.
- Partner with Storage / Image and Network teams to streamline the provisioning-to-handoff path.
- Define bring-up, rack, capacity, and acceptance standards for new GPU SKUs and data-center regions.
Job Requirement:- 3+ years (Senior: 6+ years) in large-scale bare-metal / server-fleet operations, HPC, or cloud infrastructure.
- Hands-on experience operating GPU servers at scale (e.g. NVIDIA HGX / DGX-class), including driver / CUDA and firmware management.
- Strong Linux systems skills; experience with PXE / IPMI / Redfish, OS imaging, and automated provisioning.
- Familiarity with DPU / SmartNIC (e.g. NVIDIA BlueField) and bare-metal networking.
- Infrastructure automation skills (Ansible, Terraform, Python / Go).
- Comfortable owning on-call, incident management, and operational runbooks.
- Multi-region / large-fleet operations experience is a strong plus.