AI Storage Solutions Expert

Bitdeer Technologies Group

$125K — $150K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in enterprise or HPC storage operations, with 2+ years focused on AI/ML workloads
  • Hands-on experience with WEKA, VAST Data, Ceph, or DDN/Lustre
  • Strong understanding of AI training I/O patterns including checkpointing and dataset management
  • Familiarity with high-performance storage networking technologies like NFS over RDMA and NVMe-oF
  • Knowledge of GPU Direct Storage and RDMA data transfer methods
  • Expertise in storage performance monitoring tools such as fio, IOR, and mdtest
  • Experience with implementing multi-tenant storage solutions with Quality of Service (QoS) and access controls
  • Strong grasp of Linux system internals, including kernel tuning and filesystem management
  • Innovative approach to turning operational challenges into ML-based signal detection

Responsibilities

  • Deploy and operate parallel and distributed storage systems optimized for AI workloads
  • Implement and manage multi-tenant storage isolation with QoS, quotas, and access controls
  • Configure and optimize GPU Direct Storage for efficient data paths between GPUs and storage
  • Manage storage networking and orchestration across high-speed storage fabrics
  • Analyze and enhance storage performance using profiling tools to troubleshoot and optimize operations
  • Plan and allocate storage capacity to support GPU cluster growth and customer demands
  • Instrument storage telemetry to inform predictive maintenance and incident response
  • Collaborate with platform teams to develop storage-fault prediction mechanisms and automation strategies

Benefits

  • Opportunity to shape the storage architecture of a cutting-edge AI-operated GPU cloud
  • Play a key role in preventing costly training downtime through innovative storage solutions
  • Work with advanced technologies in high-performance computing and storage
  • Engage in a collaborative environment with skilled professionals and technical teams
  • Opportunities for continuous learning and professional development in a rapidly evolving field
Full Job Description
Position Overview

You own the IO layer that trains the models - and the signals we need to predict storage faults before a checkpoint stalls a $50M training run.

Bitdeer is building an AI-operated GPU cloud. Storage is where AI workloads either fly or fall over: a slow parallel read can starve a 1,000-GPU job; a stalled checkpoint can waste a full training epoch. In this role you deploy and operate the high-performance storage layer for AI training and inference across NeoCloud's US DCs, and you feed the AIOps substrate with the signals it needs to catch storage regressions before they page a customer.

What you'll own
  • Deploy and operate parallel/distributed storage systems: WEKA, VAST Data, Ceph, DDN/Lustre. Design storage architectures optimized for AI workload patterns - checkpoint I/O bursts, sequential dataset reads, KV cache for inference.
  • Implement multi-tenant storage isolation with per-tenant QoS, quotas, and access controls; configure and optimize GPU Direct Storage for direct GPU-to-storage data paths.
  • Deploy and manage storage networking (NFS over RDMA, NVMe-oF, high-speed storage fabrics) and Nvidia CMX for cluster-wide storage orchestration.
  • Diagnose and tune storage performance: IOPS, throughput, latency profiling with fio, IOR, mdtest; own the runbook for common failure modes.
  • Plan storage capacity aligned with GPU cluster growth and customer workload projections; manage firmware, data migration, and DR procedures.

Feed the AIOps substrate
  • Instrument storage telemetry - IO tail latency, checkpoint durations, NVMe SMART, filesystem health, RDMA counters - into the metrics/logs/traces store the platform team runs.
  • Partner with the platform team to define the storage-fault predictor: which signals, which labels (from your incidents), which false-positive tolerances.
  • Convert every novel incident into an automation: SOPs become runbook-as-code, runbook-as-code becomes an agent-executable remediation.

What success looks like in year 1
  • Observability and a baseline predictor for the top 3 storage-fault classes on our fabric.
  • Storage-incident MTTR measurably lower than at hire.
  • The Nvidia GB200-class clusters we build out ship on your storage design.

Job Requirement:
  • 5+ years in enterprise or HPC storage operations, with at least 2 years supporting AI/ML workloads
  • Hands-on deployment and operations experience with at least two of: WEKA, VAST Data, Ceph, DDN/Lustre
  • Strong understanding of AI training I/O patterns: checkpoint frequency, dataset loading, shuffle buffers
  • Experience with high-performance storage networking (NFS over RDMA, NVMe-oF)
  • Knowledge of GPU Direct Storage and RDMA-based data transfer
  • Proficiency in storage performance benchmarking and tuning (fio, IOR, mdtest)
  • Experience implementing multi-tenant storage with isolation and QoS
  • Strong Linux systems knowledge (kernel tuning, filesystem internals, block device management)
  • Instinct for turning ops toil into ML signal - you've either shipped an anomaly detector for storage/IO telemetry or you can articulate the labels and features you'd need to.
  • Runbook-as-code mindset - every SOP you write should be executable by a machine within a quarter.

Similar Jobs

More Jobs at Bitdeer Technologies Group

More Enterprise Technology Jobs

Find similar AI Storage Solutions Expert jobs: