Full Job Description
We9re looking for stellar on-policy RL engineers to work on creating robust robot policies.
We9re still a small team-which means high ownership, high equity, and the chance to shape the product from the ground up.
VLAs and other great 4base policies4 for robotics achieve ~80% success rates, but in real robot deployments, it9s essential to achieve 99.99% success rates. We can9t ask our customers to tolerate our robots packing three socks into a bin instead of four, or swapping shipping labels between two packages-not even once!
You should apply for this role if:
- You9ve re-implemented core RL algorithms (SAC, DDPG) from scratch and can debug unstable gradients / tune hyperparameters correctly
- You9ve made meaningful intellectual contributions to sample-efficient RL algorithms (e.g., DreamerV3 and MuZero)
- You9ve shipped on-policy RL on hardware that learns in the real-world (e.g., for quadruped walking or drone racing)
You9d be joining a company that already has a solid core business-with working hardware, delighted customers, and profitable unit economics. Reflex is de-risked enough to see the hazy outlines of success, but still small enough that there9s enormous upside up for grabs.
This is a rare opportunity to help build a flagship robotics company from the ground up-and to do work that will truly matter, reshaping what people believe is possible in robotics.
We love to see the things you9ve worked on. Have a portfolio or insane project you9ve worked on? Share it. We9re looking for people who push past the status quo, are passionate at work and in their own time-we9re looking for people who want to win.