ML Researcher - Posttraining

Krea

$150K — $180K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Proven experience fine-tuning diffusion models for image or video generation.
  • Expertise in large-scale model training, inference, and optimization.
  • Strong grasp of LLM and diffusion post-training pipelines and algorithms.
  • Proficiency in PyTorch and its inner workings.
  • Background in distributed training paradigms and parallelism strategies.
  • Knowledge of low precision training and inference techniques.
  • Insight into efficient inference engines and existing RL frameworks.

Responsibilities

  • Finetune diffusion models at scale for enhanced image quality and aesthetics.
  • Implement advanced post-training techniques such as supervised fine-tuning and reinforcement learning.
  • Create evaluation suites and reward designs for reinforcement learning in image generation.
  • Train reward models using custom vision-language models as part of reward design.
  • Develop and fine-tune large language models for prompt expansion using reinforcement learning.
  • Collaborate with data teams to manage data collection and model evaluations.
  • Ensure safety alignment for open-source model releases.

Benefits

  • Collaborate with a leading team in AI creative tooling.
  • Broad scope for impactful contributions within the company.
  • Comprehensive health and dental insurance coverage for employees.
  • Flexible paid time off to promote work-life balance.
  • 401k plan with a 4% company match for retirement savings.
  • On-site meals provided for breakfast, lunch, and dinner.
  • Transportation support with Uber rides covered to and from the office.
Full Job Description
We're looking for someone with deep experience in large-scale finetuning of diffusion models and posttraining techniques. We have loads of high-quality data and aim to enhance the quality and aesthetics of the models we're building.

What you'll do
  • Finetune diffusion models at scale to improve image aesthetics and quality.
  • Implement posttraining techniques ranging from supervised finetuning, preference optimization, reinforcement learning, on-policy distillation, and various distillation / acceleration techniques.
  • Design comprehensive eval suites and reward designs for the reinforcement learning stage focused on image space.
  • Train custom VLM as reward models as part of our reward design.
  • Train custom LLMs for prompt expansion through finetuning and reinforcement learning.
  • Coordinate with data teams and partners to manage collection of preference data and model evaluation results.
  • Work on safety alignment of our models for open source release.
  • Collaborate with our AI research and engineering teams to integrate advancements into our products.
What we're looking for
  • Proven work of posttraining diffusion models for image or video generation.
  • Experience with large-scale model training, inference, and optimization.
  • Strong understanding of both LLM and diffusion post training pipelines and algorithms such as PPO, GRPO, DPO, OPD, and MOPD.
  • Strong proficiency in PyTorch and understanding of its inner workings.
  • Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP. Knowing how different parallelism strategies work together and their tradeoffs.
  • Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8.
  • Good understanding of algorithms and techniques used in fast inference engines such as vLLM and sglang as well as existing RL frameworks in LLM space such as slime, miles, tinker, and verl.
  • Understanding of various RL infrastructure and optimization techniques such as async RL, fast weight transfer, pipelining rollouts, managing off policy data.
  • Ability to monitor model regression and identify weak areas and turn them into concrete evals and reward design.
  • Experience training VLM models. Many of our custom reward models use VLM to provide reward signals for our models.
  • Keeping up with the developments in related fields such as LLM, VLM, representation learning, and robotics research.
  • Being comfortable working in a goal-oriented research environment.
  • Having good judgement around when one should explore different training strategies and when it's time to commit to a specific strategy to scale compute and data.
  • Comfortable working with underspecified goals. We expect every technical member to take an ambiguous research goal and break it down into concrete requirements, plans, experiment plan, and execution items.
  • Good research taste - bias towards simplicity and methods that scale well with compute, data, and minimal human supervision.
What we offer
  • Team: Work alongside a world-class team building the future of AI creative tooling
  • Impact: Significant scope and company-wide impact
  • Competitive compensation: generous salary & equity packages
  • Health & wellness: 100% health & 99% dental/vision insurance premiums covered for employees, health FSA accounts, & long-term disability coverage
  • Time off: Flexible PTO policy
  • Financial planning: 401k with a 4% company-sponsored match
  • Meals in the office: breakfast, lunch, dinner - you name it, we'll cover it
  • Transit: Ubers covered to & from the office
  • Sponsorship: We're open to sponsoring international visas where we can (e.g., STEM OPT, OPT, H-1B, O-1, E-3).
  • And more!

Please note the above benefits & perks are for full-time employees

Similar Jobs

More Jobs at Krea

  • Product Engineer
    $125K — $150K *
    San Francisco, CA 94112 (San Francisco County)
    Consumer Technology
    In-Person
  • Software Engineer, Product
    $135K — $160K *
    San Francisco, CA 94112 (San Francisco County)
    Information Technology
    In-Person

More Information Technology Jobs

Find similar ML Researcher - Posttraining jobs: