GPU Compute & Bare Metal / DPU Engineer

Bitdeer Technologies Group

$130K — $155K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years in large-scale bare-metal/server fleet operations, HPC, or cloud infrastructure (6+ years for Senior)
  • Hands-on experience with GPU servers at scale, including driver and firmware management
  • Strong skills in Linux systems and familiarity with OS imaging/automated provisioning
  • Experience with DPU/SmartNIC technologies
  • Infrastructure automation skills using tools like Ansible, Terraform, and programming in Python/Go
  • Experience in incident management and ownership of operational runbooks
  • Multi-region operations experience is a strong plus.

Responsibilities

  • Own the full lifecycle of bare-metal GPU nodes from provisioning to decommissioning in multiple regions
  • Build automated delivery pipelines to manage fleet growth towards 10,000+ GPUs
  • Manage firmware versions and validation for DPUs/SmartNICs and servers
  • Drive improvements in fleet reliability and incident response
  • Participate in a 7x24 on-call rotation; create runbooks and tools to minimize manual work
  • Collaborate with Storage/Image and Network teams to streamline processes
  • Define standards for new GPU SKUs and data center regions.

Benefits

  • Opportunity to work with cutting-edge GPU technology
  • Join a scalable, growing team with a clear pathway to manage a large fleet
  • Gain expertise in multi-region fleet operations
  • Access to automation tools and practices
  • Participation in a sustainable on-call rotation to enhance work-life balance.
Full Job Description
Position Overview

As the platform scales toward a 10,000+ GPU, multi-region footprint, we are expanding the team that turns procured GPU capacity into reliable, sellable compute. You will own the end-to-end lifecycle of bare-metal GPU nodes-from provisioning and delivery through break-fix, firmware/DPU management, and decommissioning-directly driving delivery throughput, fleet availability, and the ROI of our largest capital investment.

Key Responsibilities
  • Own the full lifecycle of bare-metal GPU nodes: provisioning, delivery / onboarding, in-service operation, break-fix, and decommissioning across multiple regions.
  • Build and operate automated, repeatable node-delivery pipelines to eliminate the current delivery backlog and keep pace with fleet growth toward 10,000+ GPUs.
  • Manage DPU / SmartNIC and server firmware (BMC / BIOS / NIC / GPU firmware): version baselines, upgrades, and validation.
  • Drive fleet reliability: reduce MTTR, lead incident response and root-cause analysis, and improve hardware-health monitoring.
  • Participate in a sustainable 7x24 multi-region on-call rotation; build runbooks and tooling that reduce manual toil.
  • Partner with Storage / Image and Network teams to streamline the provisioning-to-handoff path.
  • Define bring-up, rack, capacity, and acceptance standards for new GPU SKUs and data-center regions.

Job Requirement:
  • 3+ years (Senior: 6+ years) in large-scale bare-metal / server-fleet operations, HPC, or cloud infrastructure.
  • Hands-on experience operating GPU servers at scale (e.g. NVIDIA HGX / DGX-class), including driver / CUDA and firmware management.
  • Strong Linux systems skills; experience with PXE / IPMI / Redfish, OS imaging, and automated provisioning.
  • Familiarity with DPU / SmartNIC (e.g. NVIDIA BlueField) and bare-metal networking.
  • Infrastructure automation skills (Ansible, Terraform, Python / Go).
  • Comfortable owning on-call, incident management, and operational runbooks.
  • Multi-region / large-fleet operations experience is a strong plus.


Similar Jobs

More Jobs at Bitdeer Technologies Group

More Information Technology Jobs

Find similar GPU Compute & Bare Metal / DPU Engineer jobs: