Role ImpactYou'll build and operate the storage systems that feed frontier AI workloads. Own reliable, high-throughput access to datasets, checkpoints, and model artifacts, balancing performance, durability, availability, and cost as GPU clusters scale.
Core Technical Responsibilities- Design and operate storage architectures for training datasets, checkpointing, inference artifacts, and shared research workflows
- Deploy and tune parallel filesystems, object storage, and local NVMe caching for demanding AI workloads
- Benchmark throughput, latency, metadata performance, and concurrent access with representative training and checkpoint workloads
- Build provisioning, capacity planning, lifecycle management, and operational automation for storage services
- Design and test replication, recovery, backup, and failure-handling procedures with explicit durability and availability targets
- Diagnose performance and reliability issues across applications, clients, networks, filesystems, and devices
- Implement access controls, tenant separation, quotas, monitoring, and runbooks; collaborate with compute and networking teams
Technical RequirementsRequired Experience- 3+ years building or operating production distributed storage systems
- Hands-on experience with at least one parallel or distributed filesystem or object storage platform, such as Lustre, BeeGFS, Ceph, or GPFS
- Strong Linux administration and performance troubleshooting skills
- Experience automating infrastructure operations in Python, Go, Bash, or similar languages
- Understanding of storage failure modes, data integrity, consistency, replication, and recovery
Infrastructure Skills- Block, file, and object storage semantics and their performance tradeoffs
- NVMe/SSD performance, filesystem tuning, I/O profiling, and benchmarking
- High-throughput storage networking and distributed client behavior
- Capacity forecasting, observability, alerting, and safe maintenance procedures
- Authentication, authorization, encryption, and secure data lifecycle management
Nice to Have- Experience supporting large GPU training clusters and high-volume checkpoint workloads
- S3-compatible object storage, data tiering, or distributed caching
- RDMA-enabled storage or GPUDirect Storage experience
- Kubernetes storage integrations or SLURM environments
- Storage cost optimization and contributions to open-source storage systems
Growth OpportunityYou'll work directly with customers pushing the boundaries of AI, from startups training foundation models to enterprises deploying massive inference infrastructure. You'll collaborate with our world-class engineering team while having direct impact on systems powering the next generation of AI breakthroughs.
We value expertise and customer obsession - if you're passionate about building reliable, high-performance GPU infrastructure and have a track record of successful large-scale deployments, we want to talk to you.
Apply now and join us in our mission to democratize access to planetary scale computing.
CompensationCash compensation range of $150,000-$300,000 plus equity incentives.