Research Member of Technical Staff- Video Generation Modeling

Rhoda AI

$130K — $180K *
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Strong background in large-scale generative modeling (video generation or language model pretraining)
  • Hands-on experience training large generative models from scratch at scale
  • Deep understanding of autoregressive modeling and scaling behavior
  • Fluency with modern ML frameworks (especially PyTorch)
  • Ability to design experiments and interpret results effectively
  • Strong research taste and ability to identify high-leverage questions
  • Comfort in a fast-paced, ambiguous startup environment

Responsibilities

  • Design and train large-scale causal video generation models on web-scale data
  • Develop and validate training objectives and model architectures for video prediction
  • Research scaling laws and data efficiency for web-scale video pretraining
  • Investigate video properties that enhance robotic control and action prediction
  • Build evaluations to measure video generation quality and prediction fidelity
  • Conduct ablations and benchmarking to understand model quality
  • Collaborate with various teams to translate research into working systems
  • Publish and present research at top-tier ML and robotics venues

Benefits

  • Work on innovative approaches to robot learning
  • Enable robots to understand and predict visual environments
  • Collaborate closely across teams with no silos
  • Experience high ownership and fast iteration in a small team
Full Job Description
Were looking for Research Scientists and Research Engineers to push the frontier of large-scale pre-training for our video action model. Our approach formulates robot control as video prediction - we pre-train causal video generation models on web-scale video data, then adapt them to predict robot actions from real-world demonstrations. Youll work on the core architectures, training objectives, and scaling strategies that determine how well our models learn from internet-scale video. We hire across levels - from senior to staff - and welcome both research-track and engineering-track candidates.

What Youll Do
  • Design and train large-scale causal video generation models on web-scale video data
  • Develop and validate training objectives, model architectures, and data mixtures for video prediction at scale
  • Research scaling laws and data efficiency for web-scale video pretraining
  • Investigate what properties of web video transfer most effectively to robotic control and action prediction
  • Build systematic evaluations to measure video generation quality, long-horizon prediction fidelity, and downstream robot task performance
  • Run rigorous ablations and benchmarking to understand what drives model quality at scale
  • Collaborate closely with data & evaluation, post-training, and training systems teams to translate research ideas into working systems
  • Publish and present work at top-tier ML and robotics venues (especially valued for RS track)

What Were Looking For
  • Strong background in large-scale generative modeling - either video generation (autoregressive video models, diffusion transformers, causal video architectures) or language model pretraining (LLMs, autoregressive transformers at scale)
  • Hands-on experience training large generative models from scratch at scale
  • Deep understanding of autoregressive modeling, causal architectures, and scaling behavior
  • Fluency with modern ML frameworks (PyTorch required; JAX a plus)
  • Ability to design experiments, interpret results, and iterate quickly
  • Strong research taste: ability to identify high-leverage questions and cut through noise
  • Comfort operating in a fast-moving, ambiguous startup environment
  • Staff-level candidates are expected to define technical direction and drive research strategy independently; senior/MTS candidates execute complex projects with strong fundamentals and growing scope

Nice to Have (But Not Required)
  • PhD in ML, CS, Robotics, or a related field - or equivalent research/industry experience
  • Strong publication record at NeurIPS, ICML, ICLR, CVPR, CoRL, etc. (especially valued for RS track)
  • Prior work specifically on video generation models (autoregressive video, diffusion transformers, world models, or causal video architectures)
  • Experience with large-scale autoregressive language model pretraining and scaling
  • Familiarity with web-scale video datasets and video data curation pipelines
  • Prior work connecting video generation to control, action prediction, or robotic learning
  • Familiarity with distributed training and multi-node infrastructure

Why This Role
  • Work on a fundamentally different approach to robot learning - web-scale video pretraining rather than robot-data-only VLA models
  • Your models give our robots the ability to understand and predict the visual world from internet-scale supervision
  • Direct collaboration with data, post-training, and deployment teams with no silos
  • High ownership and fast iteration in a small, elite team

Similar Jobs

More Jobs at Rhoda AI

More Consumer Technology Jobs

Find similar Research Member of Technical Staff- Video Generation Modeling jobs: