Reinforcement Learning Infrastructure Engineer

Elorian

$275K — $475K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of experience in distributed systems and large-scale reinforcement learning training pipelines
  • Experience with actor-learner architectures and scalable environment rollouts
  • Strong proficiency in Python, with skills in PyTorch or JAX
  • Familiarity with async training infrastructure and simulation-based frameworks
  • Experience with multi-node GPU orchestration using tools like Ray, SLURM, or Kubernetes
  • Proven ability to enhance training throughput and GPU utilization
  • Strong engineering skills for writing maintainable code and debugging complex systems

Responsibilities

  • Design and optimize infrastructure for large-scale RL and post-training workloads
  • Enhance reliability, scalability, and throughput of distributed RL training pipelines
  • Build actor-learner architectures and manage environment rollouts on a large scale
  • Develop monitoring and observability tools ensuring uptime and reproducibility of RL systems
  • Collaborate with researchers to implement algorithmic concepts into practical training pipelines
  • Improve GPU utilization across the training cluster

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support as needed
Full Job Description
The Role

We're looking for an infrastructure engineer to design and build the core systems behind how we train our models with reinforcement learning (RL).

You'll own the training infrastructure end to end, from rollout and reward pipelines to orchestration, reliability, and observability. The work spans both the algorithmic side of RL and the systems reality of running distributed training at scale, and you'll partner closely with our research team to keep RL training fast, stable, and dependable for the multimodal, visual reasoning models at the center of our work.

What You Will Do
  • Design, build, and optimize the infrastructure that powers our large-scale RL and post-training workloads
  • Improve the reliability, scalability, and throughput of distributed RL training pipelines
  • Build actor-learner architectures and orchestrate environment rollouts at scale
  • Develop monitoring and observability tools that ensure high uptime, debuggability, and reproducibility across RL systems
  • Collaborate with researchers to translate algorithmic ideas into production-grade training pipelines
  • Improve GPU utilization and training throughput across the cluster


What We're Looking For

Minimum qualifications:
  • 3+ years of distributed systems experience, including building or optimizing large-scale RL training pipelines (PPO, GRPO, or similar on-policy methods)
  • Experience with actor-learner architectures and environment rollout orchestration at scale
  • Strong Python skills, plus PyTorch or JAX
  • Experience with async training infrastructure, replay buffers, or simulation-based environment frameworks
  • Multi-node GPU orchestration experience (Ray, SLURM, or Kubernetes)
  • A track record of improving training throughput and GPU utilization at scale
  • Strong engineering skills; ability to contribute performant, maintainable code and debug in complex codebases

Preferred qualifications (strong candidates may have some, not all):
  • Experience with multimodal or agentic RL environments
  • Experience with RLHF or reward modeling pipelines
  • A self-directed builder who moves quickly and works across teams in an early-stage setting


Logistics

Location: This role is based on-site in Palo Alto, California.

Compensation: Depending on background, skills, and experience, the expected annual base salary range for this position is $275,000 - $475,000 USD, plus equity and benefits.

Visa sponsorship: We sponsor work visas. We can't promise every case will succeed, but for the right person we'll work through the process with you.

Benefits: We offer health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Similar Jobs

More Jobs at Elorian

More Information Technology Jobs

Find similar Reinforcement Learning Infrastructure Engineer jobs: