Position OverviewWe are seeking a Senior AI Storage Infrastructure Engineer to build the critical data-delivery fabric of our AI-native NeoCloud. AI model training and inference are profoundly I/O intensive; you will be responsible for architecting high-performance storage solutions that eliminate bottlenecks and ensure GPUs are constantly saturated with data. This role sits at the intersection of distributed storage, kernel-level I/O, and Kubernetes orchestration. You will design the pathways-from NVMe-backed local caching for massive LLM weights to parallel file system integration-that enable seamless, low-latency access for large-scale distributed training and inference workloads.
Key Responsibilities- Design, deploy, and maintain robust Container Storage Interface (CSI) drivers for high-performance parallel file systems (e.g., Weka, Lustre, DAOS, VAST).
- Architect and implement GPUDirect Storage (GDS) integrations to enable direct memory access (DMA) between NVMe drives and GPU memory, bypassing CPU bottlenecks.
- Develop and manage local NVMe caching strategies for rapid, low-latency loading of massive model weights and datasets during distributed training.
- Optimize IOPS, throughput, and latency profiles across the entire containerized storage stack, from the storage array to the container runtime.
- Collaborate with the GPU Systems & Fabric team to ensure the storage layer is fully optimized for RDMA and high-speed interconnects (InfiniBand, RoCE).
- Implement automated monitoring and alerting for storage performance, detecting and mitigating I/O contention or hardware degradation before it impacts production jobs.
- Define storage policies, quota management, and multi-tenancy isolation strategies within Kubernetes to ensure fair resource sharing for customer workloads.
- Mentor junior engineers and drive architectural design reviews to maintain high standards of reliability and performance across the infrastructure team.
Qualifications- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field.
- 5+ years of experience in distributed storage systems and high-performance file systems, with a deep understanding of POSIX compliance and file I/O semantics.
- Deep expertise in the Kubernetes CSI paradigm, including building or extending volume plugins and storage operators.
- Strong hands-on experience with block/file I/O at the Linux OS level and kernel-level performance tuning.
- Familiarity with high-throughput networking protocols (RDMA, InfiniBand, RoCE) and how they interact with storage subsystems.
- Proven track record of operating, debugging, and scaling large-scale storage environments in production or HPC settings.
- Experience with infrastructure automation tools (e.g., Terraform, Ansible) and CI/CD pipelines.
- Excellent technical communication skills, with the ability to influence cross-functional architectural decisions.
- Experience working in high-velocity, high-growth engineering environments is strongly preferred.