What you'll do- Train diffusion models for image and video generation on large GPU clusters.
- Fully optimize and profile large distributed training runs across model architectures, kernels, data loading, memory constraints, and communication.
- Implement and improve various distributed training strategies including FSDP, CP, SP, TP, and EP.
- Continuously improve model quality and reliability through data, model architecture, training pipeline, structuring experiments, and eval design.
- Debug distributed training errors and implement fault tolerance solutions, identifying bad GPU, NVLink, Infiniband (IB) components as well as monitoring numerical errors and NCCL issues.
- Ablate different architecture, attention, optimizer, data, and algorithmic choices to reliably improve efficiency and performance of our models.
What we're looking for- Proven track record in working with image or video models at scale (publications or open-source contributions a plus).
- Strong proficiency in PyTorch and understanding of its inner workings.
- Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP. Knowing how different parallelism strategies work together and their tradeoffs.
- Experience in profiling and debugging large distributed training. Being comfortable with analyzing traces to identify bottlenecks and look for improvements.
- Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8.
- Solid understanding of diffusion model training pipeline across pretraining, midtraining, preference optimization, and reinforcement learning.
- Keeping up with the developments in related fields such as LLM, VLM, representation learning, and robotics research.
- Being comfortable working in a goal-oriented research environment.
- Having good judgement around when one should explore different training strategies and when it's time to commit to a specific strategy to scale compute and data.
- Comfortable working with underspecified goals. We expect every technical member to take an ambiguous research goal and break it down into concrete requirements, plans, experiment plan, and execution items.
- Good research taste - bias towards simplicity and methods that scale well with compute, data, and minimal human supervision.
- Ability to iterate rapidly, and propose creative research directions.
- Be comfortable getting your hands dirty with data and designing custom data pipelines to improve data quality.
What we offer- Team: Work alongside a world-class team building the future of AI creative tooling
- Impact: Significant scope and company-wide impact
- Competitive compensation: generous salary & equity packages
- Health & wellness: 100% health & 99% dental/vision insurance premiums covered for employees, health FSA accounts, & long-term disability coverage
- Time off: Flexible PTO policy
- Financial planning: 401k with a 4% company-sponsored match
- Meals in the office: breakfast, lunch, dinner - you name it, we'll cover it
- Transit: Ubers covered to & from the office
- Sponsorship: We're open to sponsoring international visas where we can (e.g., STEM OPT, OPT, H-1B, O-1, E-3).
- And more!
Please note the above benefits & perks are for full-time employees