Position OverviewGPU training and inference at 10,000+ GPU, multi-region scale depend on high-throughput, low-latency storage that can sustain massive parallel I/O. We are looking for an engineer who deeply understands distributed file systems and can integrate distributed / parallel storage systems into our GPU cloud - covering performance, multi-tenancy, and reliability - while also owning the image / driver / registry pipeline on the node-delivery critical path.
Key Responsibilities- Design and integrate distributed / parallel file systems (e.g. Ceph, Lustre, GPFS / Spectrum Scale, BeeGFS, JuiceFS) into the GPU cloud, optimized for AI training / inference I/O patterns.
- Own end-to-end distributed-storage integration: provisioning, mounting, multi-tenant isolation, quota, and lifecycle within the platform / control plane.
- Tune storage throughput and latency for large-scale parallel access (dataset loading, checkpointing); benchmark across GPU SKUs and workloads.
- Architect multi-region storage: data locality, replication / consistency, durability (failure domains), and cross-region access.
- Own golden images, templates, GPU drivers / CUDA, and the container / image registry, including versioned release and multi-region distribution. (Secondary scope.)
- Build monitoring, capacity planning, and runbooks; eliminate single points of failure.
- Partner with Compute (delivery), Network (storage fabric / RDMA), and Control Plane (provisioning / quota) teams.
Job Requirement:- 3+ years (Senior 6+) in storage engineering or platform infrastructure, with hands-on distributed / parallel file system experience.
- Strong understanding of distributed file system internals - data / metadata separation, replication, consistency models, POSIX vs object semantics.
- Proven experience integrating and operating distributed storage in production (e.g. Ceph, Lustre, GPFS / Spectrum Scale, BeeGFS, JuiceFS, MinIO).
- Performance tuning for high-throughput / parallel I/O; familiarity with NVMe, RDMA / RoCE storage networking, and caching is a strong plus.
- Strong Linux systems depth and automation skills (Python / Go, CI / CD).
- HPC / AI storage or multi-region storage experience a strong plus.