ML Researcher - Image / Video Diffusion

Krea

$150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience training image or video models at scale
  • Proficient in PyTorch with solid understanding of its internals
  • In-depth knowledge of distributed training paradigms (FSDP, CP, SP, etc.)
  • Experience in profiling and debugging large distributed training systems
  • Familiar with low precision training techniques (FP8, NVFP4)
  • Understanding of diffusion model training processes
  • Ability to navigate an ambiguous research environment

Responsibilities

  • Train diffusion models for image and video generation on large GPU clusters.
  • Optimize and profile distributed training runs for efficiency.
  • Implement and enhance various distributed training strategies.
  • Improve model quality and reliability through structured experimentation.
  • Debug distributed training errors and implement fault tolerance solutions.
  • Experiment with architecture and algorithmic choices to boost model performance.

Benefits

  • Work alongside a world-class AI development team
  • Significant impact on company-wide projects
  • Comprehensive health and wellness coverage
  • Flexible PTO policy to support work-life balance
  • 401k plan with a generous company match
  • Meal provisions for all three meals in the office
  • Covered transit expenses for commuting to the office
  • Openness to visa sponsorship for qualified candidates
Full Job Description
What you'll do
  • Train diffusion models for image and video generation on large GPU clusters.
  • Fully optimize and profile large distributed training runs across model architectures, kernels, data loading, memory constraints, and communication.
  • Implement and improve various distributed training strategies including FSDP, CP, SP, TP, and EP.
  • Continuously improve model quality and reliability through data, model architecture, training pipeline, structuring experiments, and eval design.
  • Debug distributed training errors and implement fault tolerance solutions, identifying bad GPU, NVLink, Infiniband (IB) components as well as monitoring numerical errors and NCCL issues.
  • Ablate different architecture, attention, optimizer, data, and algorithmic choices to reliably improve efficiency and performance of our models.


What we're looking for
  • Proven track record in working with image or video models at scale (publications or open-source contributions a plus).
  • Strong proficiency in PyTorch and understanding of its inner workings.
  • Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP. Knowing how different parallelism strategies work together and their tradeoffs.
  • Experience in profiling and debugging large distributed training. Being comfortable with analyzing traces to identify bottlenecks and look for improvements.
  • Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8.
  • Solid understanding of diffusion model training pipeline across pretraining, midtraining, preference optimization, and reinforcement learning.
  • Keeping up with the developments in related fields such as LLM, VLM, representation learning, and robotics research.
  • Being comfortable working in a goal-oriented research environment.
  • Having good judgement around when one should explore different training strategies and when it's time to commit to a specific strategy to scale compute and data.
  • Comfortable working with underspecified goals. We expect every technical member to take an ambiguous research goal and break it down into concrete requirements, plans, experiment plan, and execution items.
  • Good research taste - bias towards simplicity and methods that scale well with compute, data, and minimal human supervision.
  • Ability to iterate rapidly, and propose creative research directions.
  • Be comfortable getting your hands dirty with data and designing custom data pipelines to improve data quality.
What we offer
  • Team: Work alongside a world-class team building the future of AI creative tooling
  • Impact: Significant scope and company-wide impact
  • Competitive compensation: generous salary & equity packages
  • Health & wellness: 100% health & 99% dental/vision insurance premiums covered for employees, health FSA accounts, & long-term disability coverage
  • Time off: Flexible PTO policy
  • Financial planning: 401k with a 4% company-sponsored match
  • Meals in the office: breakfast, lunch, dinner - you name it, we'll cover it
  • Transit: Ubers covered to & from the office
  • Sponsorship: We're open to sponsoring international visas where we can (e.g., STEM OPT, OPT, H-1B, O-1, E-3).
  • And more!

Please note the above benefits & perks are for full-time employees

Similar Jobs

More Jobs at Krea

  • Product Engineer
    $125K — $150K *
    San Francisco, CA 94112 (San Francisco County)
    Consumer Technology
    In-Person
  • Software Engineer, Product
    $135K — $160K *
    San Francisco, CA 94112 (San Francisco County)
    Information Technology
    In-Person

More Information Technology Jobs

Find similar ML Researcher - Image / Video Diffusion jobs: