GPU Compute & Bare Metal / DPU Engineer

Bitdeer Technologies Group

$120K — $145K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years in server-fleet operations, HPC, or cloud infrastructure; 6+ years for Senior role
  • Hands-on experience with GPU servers at scale, including driver management
  • Strong Linux systems skills with experience in OS imaging and automated provisioning
  • Familiarity with DPU / SmartNIC technologies like NVIDIA BlueField
  • Infrastructure automation skills with tools like Ansible and Terraform
  • Comfortable with incident management and on-call responsibilities
  • Multi-region operational experience is preferred

Responsibilities

  • Own the full lifecycle of bare-metal GPU nodes from provisioning to decommissioning
  • Build automated delivery pipelines to manage GPU fleet scaling
  • Manage firmware for servers and SmartNICs, ensuring versions are current and validated
  • Improve fleet reliability by reducing downtime and leading incident analysis
  • Participate in a multi-region on-call rotation, creating runbooks to minimize manual work
  • Collaborate with Storage and Network teams for efficient provisioning
  • Establish standards for new GPU hardware and data center setup

Benefits

  • Flexible schedule with remote work options
  • Comprehensive health, dental, and vision insurance
  • Generous paid time off and holiday leave
  • Retirement savings plan with company matching
  • Opportunities for professional development and training
Full Job Description
Position Overview

As the platform scales toward a 10,000+ GPU, multi-region footprint, we are expanding the team that turns procured GPU capacity into reliable, sellable compute. You will own the end-to-end lifecycle of bare-metal GPU nodes-from provisioning and delivery through break-fix, firmware/DPU management, and decommissioning-directly driving delivery throughput, fleet availability, and the ROI of our largest capital investment.

Key Responsibilities
  • Own the full lifecycle of bare-metal GPU nodes: provisioning, delivery / onboarding, in-service operation, break-fix, and decommissioning across multiple regions.
  • Build and operate automated, repeatable node-delivery pipelines to eliminate the current delivery backlog and keep pace with fleet growth toward 10,000+ GPUs.
  • Manage DPU / SmartNIC and server firmware (BMC / BIOS / NIC / GPU firmware): version baselines, upgrades, and validation.
  • Drive fleet reliability: reduce MTTR, lead incident response and root-cause analysis, and improve hardware-health monitoring.
  • Participate in a sustainable 7x24 multi-region on-call rotation; build runbooks and tooling that reduce manual toil.
  • Partner with Storage / Image and Network teams to streamline the provisioning-to-handoff path.
  • Define bring-up, rack, capacity, and acceptance standards for new GPU SKUs and data-center regions.

Job Requirement:
  • 3+ years (Senior: 6+ years) in large-scale bare-metal / server-fleet operations, HPC, or cloud infrastructure.
  • Hands-on experience operating GPU servers at scale (e.g. NVIDIA HGX / DGX-class), including driver / CUDA and firmware management.
  • Strong Linux systems skills; experience with PXE / IPMI / Redfish, OS imaging, and automated provisioning.
  • Familiarity with DPU / SmartNIC (e.g. NVIDIA BlueField) and bare-metal networking.
  • Infrastructure automation skills (Ansible, Terraform, Python / Go).
  • Comfortable owning on-call, incident management, and operational runbooks.
  • Multi-region / large-fleet operations experience is a strong plus.

Similar Jobs

More Jobs at Bitdeer Technologies Group

More Information Technology Jobs

Find similar GPU Compute & Bare Metal / DPU Engineer jobs: