Research Scientist, Scaling RL

Periodic Labs

• $250K — $350K *
Consumer Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in reinforcement learning (RL) and training large language models (LLMs)
  • Bachelor's degree or equivalent experience
  • Strong attention to detail and a scientific approach to problem-solving
  • Ability to design small-scale RL experiments that scale to larger training runs
  • Comfortable navigating a complex training stack for implementation and debugging

Responsibilities

  • Design experiments to analyze RL performance scaling with compute and model size
  • Develop advanced RL algorithms focusing on policy optimization and exploration
  • Create adaptive sampling and curriculum methods to adjust task difficulty
  • Investigate bias and stability issues during RL training
  • Enhance compute efficiency through hyperparameter experimentation

Benefits

  • Visa sponsorship available
  • Opportunity to work on cutting-edge scientific tasks
  • Collaborative environment focused on innovation
  • Access to large-scale training resources and experiments
  • Potential for equity in a growing company
Full Job Description
About the Role

We're training frontier models to develop deep scientific knowledge and reasoning for scientific tasks. You'll study how RL scales with training compute, develop better algorithms, and take ideas from controlled experiments to our largest runs like Periodic Neon.

What You'll Do
  • Design experiments to understand how RL performance scales with compute, model size, data, and reward quality, building on work such as ScaleRL
  • Develop better RL algorithms, spanning policy optimization, advantage estimation, exploration, and credit assignment for long-horizon RL tasks
  • Build adaptive sampling and curriculum methods that adjust task difficulty, problem selection, and the number of rollouts as models improve
  • Study bias and stability during RL training, including importance-sampling corrections and methods to tackle policy staleness and training-inference mismatch, as discussed here.
  • Improve compute efficiency across training and inference through experiments with hyperparameters, such as length penalties, rollout counts, batch sizes, and update schedules.


You Will Thrive in This Role If You Have
  • Hands-on experience training LLMs with reinforcement learning
  • Strong attention to detail and rigorous approach to answer questions scientifically.
  • Coming up with small-scale RL setups that transfers to large-scale training runs.
  • Comfort working across a complex training stack to implement, debug, and test new research ideas.


Mechanics

Minimum experience: 5+ years
Minimum education: Bachelor's degree or similar experience

Location: Menlo Park, CA

Compensation: $250,000-$350,000 base + equity

Visa sponsorship: Yes, we sponsor visas and will do everything we can to assist in this process.

Similar Jobs

More Jobs at Periodic Labs

More Consumer Technology Jobs

Find similar Research Scientist, Scaling RL jobs: