Member of Technical Staff - Distributed Training Engineer

Liquid AI

$150K — $180K *
US-AnywhereRemote in San Francisco, CA
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in distributed systems design and implementation
  • Proficient with PyTorch Distributed DDP/FSDP, DeepSpeed ZeRO, Megatron-LM TP/PP
  • Strong skills in diagnosing performance bottlenecks and failure modes
  • Understanding of hardware accelerators and networking topologies
  • Experience with data pipeline optimization for machine learning workloads

Responsibilities

  • Design and build core systems to enhance training speed and reliability
  • Create scalable distributed training infrastructure for GPU clusters
  • Implement and tune parallelism and sharding for evolving ML architectures
  • Optimize distributed efficiency, including topology-aware collectives
  • Develop robust data loading systems to eliminate I/O bottlenecks
  • Create mechanisms for recovery and memory management in training
  • Build tools for monitoring, profiling, and debugging training processes

Benefits

  • Opportunity to build systems from scratch for novel architectures
  • 100% coverage of medical, dental, and vision premiums for employees and dependents
  • 401(k) matching up to 4% of base pay
  • Unlimited PTO plus company-wide Refill Days throughout the year
Full Job Description
Our Training Infrastructure team is building the distributed systems that power our next-generation Liquid Foundation Models. As we scale, we need to design, implement, and optimize the infrastructure that enables large-scale training. This is a high-ownership training systems role focused on runtime/performance/reliability (not a general platform/SRE role). You'll work on a small team with fast feedback loops, building critical systems from the ground up rather than inheriting mature infrastructure. We need someone who: - Loves distributed systems complexity: Our team builds systems that keeps long training runs stable, debugs training failures across GPU clusters, and improves performance. - Wants to build: We need builders who find satisfaction in robust, fast, reliable infrastructure. - Thrives in ambiguity: Our systems support model architectures that are still evolving. We make decisions with incomplete information and iterate quickly. - Aligns with team priorities and delivers: Our best engineers align with team priorities while pushing back with data when they see problems. The Work - Design and build core systems that make large training runs fast and reliable - Build scalable distributed training infrastructure for GPU clusters - Implement and tune parallelism/sharding strategies for evolving architectures - Optimize distributed efficiency (topology-aware collectives, comm/compute overlap, straggler mitigation) - Build data loading systems that eliminate I/O bottlenecks for multimodal datasets - Develop checkpointing mechanisms balancing memory constraints with recovery needs - Create monitoring, profiling, and debugging tools for training stability and performance Desired Experience Must-have: - Hands-on experience building distributed training infrastructure (PyTorch Distributed DDP/FSDP, DeepSpeed ZeRO, Megatron-LM TP/PP) - Experience diagnosing performance bottlenecks and failure modes (profiling, NCCL/collectives issues, hangs, OOMs, stragglers) - Understanding of hardware accelerators and networking topologies - Experience optimizing data pipelines for ML workloads Nice-to-have: - MoE (Mixture of Experts) training experience - Large-scale distributed training (100+ GPUs) - Open-source contributions to training infrastructure projects What Success Looks Like (Year One) - Training throughput has increased - Overall training efficiency/cost has improved - Training stability has improved (fewer failures, faster recovery) - Data loading bottlenecks are eliminated for multimodal workloads What We Offer - Greenfield challenges: Build systems from scratch for novel architectures. High ownership from day one. - Compensation: Competitive base salary with equity in a unicorn-stage company - Health: We pay 100% of medical, dental, and vision premiums for employees and dependents - Financial: 401(k) matching up to 4% of base pay - Time Off: Unlimited PTO plus company-wide Refill Days throughout the year

Similar Jobs

More Jobs at Liquid AI

More Information Technology Jobs

Find similar Member of Technical Staff - Distributed Training Engineer jobs: