Member of Technical Staff - Storage Infrastructure

Prime Intellect

$150K — $300K *
US-AnywhereRemote in San Francisco, CA
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of experience building or managing distributed storage systems
  • Proficiency with parallel or distributed filesystems (e.g., Lustre, Ceph)
  • Strong Linux system administration skills
  • Experience writing automation scripts in Python, Go, or Bash
  • Knowledge of storage systems' reliability, integrity, and recovery mechanisms

Responsibilities

  • Design storage architectures for AI datasets and workflows
  • Deploy and optimize parallel filesystems and object storage
  • Benchmark performance metrics for training and checkpointing tasks
  • Manage provisioning, capacity, and automation for storage services
  • Develop and test backup and recovery solutions ensuring data durability
  • Diagnose issues impacting storage performance and reliability
  • Collaborate with teams on access controls and monitoring processes

Benefits

  • Opportunity to work directly with cutting-edge AI projects
  • Collaboration with a world-class engineering team
  • Direct impact on systems driving AI advancements
  • Engagement with diverse customers from startups to enterprises
  • Focus on expertise and customer satisfaction
Full Job Description
Role Impact

You'll build and operate the storage systems that feed frontier AI workloads. Own reliable, high-throughput access to datasets, checkpoints, and model artifacts, balancing performance, durability, availability, and cost as GPU clusters scale.
Core Technical Responsibilities
  • Design and operate storage architectures for training datasets, checkpointing, inference artifacts, and shared research workflows
  • Deploy and tune parallel filesystems, object storage, and local NVMe caching for demanding AI workloads
  • Benchmark throughput, latency, metadata performance, and concurrent access with representative training and checkpoint workloads
  • Build provisioning, capacity planning, lifecycle management, and operational automation for storage services
  • Design and test replication, recovery, backup, and failure-handling procedures with explicit durability and availability targets
  • Diagnose performance and reliability issues across applications, clients, networks, filesystems, and devices
  • Implement access controls, tenant separation, quotas, monitoring, and runbooks; collaborate with compute and networking teams
Technical Requirements
Required Experience
  • 3+ years building or operating production distributed storage systems
  • Hands-on experience with at least one parallel or distributed filesystem or object storage platform, such as Lustre, BeeGFS, Ceph, or GPFS
  • Strong Linux administration and performance troubleshooting skills
  • Experience automating infrastructure operations in Python, Go, Bash, or similar languages
  • Understanding of storage failure modes, data integrity, consistency, replication, and recovery
Infrastructure Skills
  • Block, file, and object storage semantics and their performance tradeoffs
  • NVMe/SSD performance, filesystem tuning, I/O profiling, and benchmarking
  • High-throughput storage networking and distributed client behavior
  • Capacity forecasting, observability, alerting, and safe maintenance procedures
  • Authentication, authorization, encryption, and secure data lifecycle management
Nice to Have
  • Experience supporting large GPU training clusters and high-volume checkpoint workloads
  • S3-compatible object storage, data tiering, or distributed caching
  • RDMA-enabled storage or GPUDirect Storage experience
  • Kubernetes storage integrations or SLURM environments
  • Storage cost optimization and contributions to open-source storage systems
Growth Opportunity

You'll work directly with customers pushing the boundaries of AI, from startups training foundation models to enterprises deploying massive inference infrastructure. You'll collaborate with our world-class engineering team while having direct impact on systems powering the next generation of AI breakthroughs.

We value expertise and customer obsession - if you're passionate about building reliable, high-performance GPU infrastructure and have a track record of successful large-scale deployments, we want to talk to you.

Apply now and join us in our mission to democratize access to planetary scale computing.
Compensation

Cash compensation range of $150,000-$300,000 plus equity incentives.

Similar Jobs

More Jobs at Prime Intellect

More Information Technology Jobs

Find similar Member of Technical Staff - Storage Infrastructure jobs: