A Moving Experience.What You Will Work On
- Design and operate distributed training systems for large neural networks (autoregressive, diffusion, State Space Models etc.) across GPU clusters
- Optimise multi-node, multi-GPU execution to maximize throughput and utilization
- Diagnose & resolve bottlenecks across compute, memory, and network
- Improve training stability and fault tolerance at scale
- Partner with research and applied ML teams to productionize large-model training pipelines
Core Responsibilities
- Distributed Training Infrastructure
- Build and optimize GPU cluster orchestration using:
- Ensure efficient scheduling, isolation, and fairness across training workloads
- Communication & Networking
- Optimize and debug distributed communication using:
- Minimize networking bottlenecks that dominate end-to-end training time
- Scale large-model training using:
- Own multi-node launch configurations, failure recovery, and performance tuning
- Memory & Performance Optimization
- Apply advanced memory optimization techniques:
- ZeRO (Stage 1-3) and offload strategies
- Balance compute, memory, and communication to push model size and batch scale
What Success Looks Like
- GPU utilization consistently stays high (>80-90%)
- Training scales cleanly from single node to dozens or hundreds of GPUs
- Communication overhead is minimized and predictable
- Large training jobs run stably for days or weeks without failure
- New models can be trained faster, larger, and more reliably than before
Required Experience & Skills
- Deep hands-on experience with distributed systems or ML systems
- Experience running large-scale workloads on GPU clusters
- Production experience with PyTorch distributed training
- Strong understanding of parallelism strategies (data, tensor, pipeline parallelism)
- Low-level understanding of GPU communication and networking
- Critical Technical Skills
- GPU orchestration: Slurm, Kubernetes, Ray, RunAI
- Communication libraries: NCCL, RDMA, InfiniBand, NVLink
- Training frameworks: PyTorch Distributed, Megatron-LM, DeepSpeed
- Memory optimisation: activation checkpointing, ZeRO offload techniques
Common Problems You'll Be Solving
- Many teams fail at scale because:
- GPU utilization is low despite large clusters
- Networking and communication dominate training time
- Training jobs crash or become unstable at large scale
- You will be explicitly focused on eliminating these failure modes.
Ideal Background
- This role is a strong fit for individuals who have worked as:
- Distributed Systems Engineer
- AI Infrastructure Engineer
- HPC Engineer transitioning into ML
- Experience working with large language models or foundation models is a strong plus, but deep systems expertise is valued over pure model architecture experience.
Why This Role Matters
Without robust distributed training infrastructure, progress on large models stalls. This role directly enables:
- More reliable research-to-production pipelines
You will be building the foundation that makes large-scale AI possible.