Research Scientist / Engineer - Reinforcement Learning Infrastructure

Luma AI

$187K — $395K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience post-training LLMs with reinforcement learning at scale
  • Proficient in distributed PyTorch training and parallelization strategies
  • Experience creating RL environments, reward functions, and evaluators for LLM agents
  • In-depth knowledge of RL post-training frameworks and systems tradeoffs
  • Strong background with GPU clusters, networking, and communication libraries
  • Preferred: Experience running RL training across 100+ GPUs
  • Preferred: Familiarity with containerization tools like Kubernetes or Ray

Responsibilities

  • Design and scale distributed RL post-training systems for large multimodal models
  • Optimize rollout generation integrating inference engines into the training loop
  • Create RL environments for multi-step tasks with secure, scalable executions
  • Build infrastructure for verifiable rewards and defenses against reward hacking
  • Develop monitoring and debugging tools for stable large RL runs
  • Enhance RL training efficiency through advanced scheduling and resource management
  • Collaborate with researchers to execute new post-training ideas in a production environment

Benefits

  • Health, dental, and vision insurance
  • Flexible work hours and remote work options
  • Generous parental leave policy
  • Professional development opportunities and training
  • Wellness programs and fitness reimbursements
Full Job Description
About the Role

Reinforcement learning is how our foundation models go from capable to useful - learning to reason, use tools, and act over long horizons. The RL Infrastructure team builds the systems that make this possible at scale: high-throughput distributed training that couples policy optimization with large fleets of inference workers, environments that expose models to realistic multi-step tasks, and the reward, verification, and evaluation systems that turn model behavior into learning signal.

Unlike pretraining, RL at scale is a full-loop systems problem - training, rollout generation, environment execution, and reward computation all run concurrently across thousands of GPUs and must stay fast, stable, and correct together. We are looking for engineers and scientists who have lived this problem: people who have post-trained LLMs with RL, built environments and verifiers from scratch, and debugged what happens when an asynchronous rollout pipeline meets a frontier-scale training run. You will work alongside our research team to design and operate the RL stack for our largest multimodal models.

Responsibilities

  • Design, build, and scale distributed RL post-training systems for large multimodal models - orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs
  • Build and optimize high-throughput rollout generation, including efficient integration of inference engines (e.g. vLLM, SGLang) into the training loop, weight synchronization, and asynchronous / off-policy training schemes
  • Design and implement RL environments for agentic and multi-step tasks - sandboxed code execution, tool use, computer use, and multimodal interaction - that are reproducible, hermetic, and scalable to millions of episodes
  • Build reward infrastructure: verifiable / programmatic rewards, reward model serving, LLM-as-judge pipelines, and defenses against reward hacking
  • Develop the evaluation, monitoring, and debugging tooling needed to keep large RL runs stable, diagnose convergence and throughput regressions, and understand model behavior mid-run
  • Advance RL training efficiency and stability: sequence packing for long multi-turn trajectories, KV cache reuse across rollouts, curriculum and task sampling, and resource scheduling across heterogeneous training/inference workloads
  • Collaborate closely with researchers to turn new post-training ideas (RLVR, agentic RL, long-horizon credit assignment, self-improvement loops) into production-quality training runs


Experience

  • Hands-on experience post-training LLMs with reinforcement learning (e.g. PPO / GRPO-family methods, RLHF, RLVR / RL from verifiable rewards) at meaningful scale
  • Extensive experience with distributed PyTorch training and parallelization strategies (FSDP, Tensor / Pipeline / Expert Parallel) for foundation models
  • Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents - including sandboxed execution and multi-turn tool use
  • Deep familiarity with RL post-training frameworks and their systems tradeoffs (e.g. veRL, OpenRLHF, TRL, Ray-based orchestration) and inference engines used for rollouts (vLLM, SGLang)
  • Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI), and how they behave under mixed training + inference workloads
  • (Preferred) Experience running RL training across >100 GPUs, including asynchronous or disaggregated trainer/rollout architectures
  • (Preferred) Experience with containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads
  • (Preferred) Research contributions in RL for LLMs - reasoning, agents, reward modeling, or long-horizon tasks - or open-source contributions to RL training frameworks


Compensation

The base pay range for this role is $187,500 - $395,000 per year.

Similar Jobs

More Jobs at Luma AI

More Information Technology Jobs

Find similar Research Scientist / Engineer - Reinforcement Learning Infrastructure jobs: