Member of Technical Staff - Pretraining Infra

Nuance Labs

$300K — $400K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Hands-on experience with large-scale distributed training on GPU clusters (minimum hundreds, ideally 1,000+ GPUs).
  • In-depth knowledge of distributed training mechanics including various forms of parallelism and mixed precision techniques.
  • Strong understanding of GPU communication and performance debugging techniques.
  • Practical experience with major training stacks such as Megatron, PyTorch FSDP, or DeepSpeed.
  • Familiarity with omni or multimodal training challenges, especially across different data types.
  • Solid software engineering skills and a willingness to learn new technologies.

Responsibilities

  • Own the distributed training infrastructure for omni model pretraining from design to scaling.
  • Build and maintain the core training runtime, including job orchestration and monitoring.
  • Optimize training performance across various dimensions, including memory usage and end-to-end efficiency.
  • Develop infrastructure for handling diverse data types in training workloads.
  • Adapt the platform as model architectures and training strategies evolve.

Benefits

  • Health savings account plan with substantial company contributions.
  • 15 days PTO plus public holidays, with a full week off during holidays.
  • Free lunches, snacks, and beverages provided on workdays.
  • Commuter benefits available.
  • Planned 401K options under development.
Full Job Description
About the Role

We're looking for a deeply technical MTS to own distributed training infrastructure for large-scale omni model pretraining.

This role sits at the intersection of research, systems, and GPU-scale execution - building the training stack from 01 and scaling it: distributed execution, parallelism, GPU communication, data loading, checkpointing, observability, and debugging.

Our models are omni from the ground up (audio, video, language, real-time full-duplex), which introduces systems challenges beyond standard LLM training: multimodal synchronization, long temporal context, variable sequence lengths, and tight memory/throughput constraints.

High ownership. Direct impact on what models we can train, how fast research can iterate, and how reliably we scale.
What You'll Own
  • Own the distributed training stack for omni model pretraining, from 01 system design to 110 scaling across large GPU clusters.
  • Build and operate the core training runtime: job orchestration, distributed execution, checkpointing, recovery, monitoring, and debugging for long-running training jobs.
  • Optimize large-scale training performance across parallelism strategy, GPU communication, memory usage, data throughput, MFU, step time, and end-to-end training efficiency.
  • Build infrastructure for omni training workloads: high-throughput audio/video/text data loading, temporal alignment, variable sequence handling, multimodal synchronization, and memory-efficient training.
  • Evolve the platform as model architectures, training recipes, data mixtures, sequence lengths, hardware constraints, and research directions change.
What We're Looking For
  • Hands-on experience running large-scale distributed training jobs across large GPU clusters; experience at hundreds of GPUs minimum, 1,000+ GPUs a strong plus.
  • Deep understanding of distributed training mechanics: data/tensor/pipeline/sequence parallelism, gradient communication, collectives, mixed precision, activation checkpointing, optimizer state, memory pressure, and framework-level tradeoffs.
  • Strong understanding of GPU communication and performance debugging: NCCL, all-reduce/all-gather/reduce-scatter, communication-computation overlap, topology, synchronization, stragglers, low MFU, OOMs, checkpoint bottlenecks, and data starvation.
  • Practical experience with at least one major large-scale training stack such as Megatron, PyTorch FSDP, DeepSpeed, or equivalent internal infrastructure.
  • Understanding of omni or multimodal training challenges, especially audio/video/language data, long temporal context, variable sequence lengths, modality-specific bottlenecks, and high-throughput dataloading.
  • Strong software engineering fundamentals, curiosity, and adaptability to new model architectures, training frameworks, hardware constraints, and research ideas.
Bonus Points
  • Prior 01 experience building large-scale training infrastructure or deeply modifying core training frameworks, runtimes, checkpointing, or debugging systems.
  • Experience training large omni or multimodal models involving audio, video, text, or long-context temporal data.
  • Experience with adjacent infrastructure areas such as RL/post-training, data infrastructure, synthetic data generation, evaluation, or serving.
  • Publications or substantial open-source contributions in ML systems, distributed systems, HPC, GPU performance, or training infrastructure.
Compensation

$300,000 - $400,000 base salary, plus meaningful equity. We think long-term ownership matters and structure equity accordingly.

Logistics
  • Location: In-person in Seattle, 5 days a week - we believe in the compounding value of working shoulder-to-shoulder
  • Health: HSA plan with ~$2,000 in company contributions - about 2x what most big tech companies offer
  • PTO: 15 days + public holidays, and we close for a full week over the holidays
  • Lunch, beverages, and snacks: On us, every workday - the kind of thing that makes you actually look forward to the workday
  • Commuter benefits
  • 401K: In the works


Similar Jobs

More Jobs at Nuance Labs

More Information Technology Jobs

Find similar Member of Technical Staff - Pretraining Infra jobs: