Senior AI Storage Infrastructure Engineer

Bitdeer Technologies Group

$135K — $160K *
Technical Services
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field.
  • 5+ years of experience in distributed storage systems and high-performance file systems.
  • Deep expertise in Kubernetes CSI, including building or extending volume plugins.
  • Strong hands-on experience with block/file I/O at the Linux OS level.
  • Familiarity with high-throughput networking protocols like RDMA, InfiniBand, and RoCE.
  • Proven track record in operating and scaling large-scale storage environments.
  • Experience with infrastructure automation tools like Terraform or Ansible.

Responsibilities

  • Design, deploy, and maintain robust Container Storage Interface (CSI) drivers for parallel file systems.
  • Architect and implement GPUDirect Storage (GDS) integrations for direct memory access between NVMe and GPU.
  • Develop and manage local NVMe caching strategies for low-latency loading of model weights.
  • Optimize IOPS, throughput, and latency profiles across the storage stack.
  • Collaborate with the GPU Systems team to optimize storage for high-speed interconnects.
  • Implement automated monitoring for storage performance and mitigate I/O contention issues.
  • Define storage policies and manage multi-tenancy isolation within Kubernetes.

Benefits

  • Flexible work arrangements to promote work-life balance.
  • Opportunities for professional development and continuing education.
  • Access to cutting-edge technologies and tools in AI and data infrastructure.
  • Collaborative and innovative team environment with a focus on mentorship.
Full Job Description
Position Overview

We are seeking a Senior AI Storage Infrastructure Engineer to build the critical data-delivery fabric of our AI-native NeoCloud. AI model training and inference are profoundly I/O intensive; you will be responsible for architecting high-performance storage solutions that eliminate bottlenecks and ensure GPUs are constantly saturated with data. This role sits at the intersection of distributed storage, kernel-level I/O, and Kubernetes orchestration. You will design the pathways-from NVMe-backed local caching for massive LLM weights to parallel file system integration-that enable seamless, low-latency access for large-scale distributed training and inference workloads.

Key Responsibilities
  • Design, deploy, and maintain robust Container Storage Interface (CSI) drivers for high-performance parallel file systems (e.g., Weka, Lustre, DAOS, VAST).
  • Architect and implement GPUDirect Storage (GDS) integrations to enable direct memory access (DMA) between NVMe drives and GPU memory, bypassing CPU bottlenecks.
  • Develop and manage local NVMe caching strategies for rapid, low-latency loading of massive model weights and datasets during distributed training.
  • Optimize IOPS, throughput, and latency profiles across the entire containerized storage stack, from the storage array to the container runtime.
  • Collaborate with the GPU Systems & Fabric team to ensure the storage layer is fully optimized for RDMA and high-speed interconnects (InfiniBand, RoCE).
  • Implement automated monitoring and alerting for storage performance, detecting and mitigating I/O contention or hardware degradation before it impacts production jobs.
  • Define storage policies, quota management, and multi-tenancy isolation strategies within Kubernetes to ensure fair resource sharing for customer workloads.
  • Mentor junior engineers and drive architectural design reviews to maintain high standards of reliability and performance across the infrastructure team.

Qualifications
  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field.
  • 5+ years of experience in distributed storage systems and high-performance file systems, with a deep understanding of POSIX compliance and file I/O semantics.
  • Deep expertise in the Kubernetes CSI paradigm, including building or extending volume plugins and storage operators.
  • Strong hands-on experience with block/file I/O at the Linux OS level and kernel-level performance tuning.
  • Familiarity with high-throughput networking protocols (RDMA, InfiniBand, RoCE) and how they interact with storage subsystems.
  • Proven track record of operating, debugging, and scaling large-scale storage environments in production or HPC settings.
  • Experience with infrastructure automation tools (e.g., Terraform, Ansible) and CI/CD pipelines.
  • Excellent technical communication skills, with the ability to influence cross-functional architectural decisions.
  • Experience working in high-velocity, high-growth engineering environments is strongly preferred.


Similar Jobs

More Jobs at Bitdeer Technologies Group

More Technical Services Jobs

Find similar Senior AI Storage Infrastructure Engineer jobs: