About the roleWe are building a recursive self-improvement system - a machine learning system that iteratively improves itself through feedback, evaluation, and automated learning loops. You will help build the engineering pipeline that keeps these loops fast, reliable, and trustworthy: the training and evaluation pipelines, the reward and feedback signals, and the safeguards that prevent a self-improving system from silently degrading or gaming its objectives.
This is an engineering-first role with deep reinforcement learning requirements. You should be equally comfortable writing robust production ML code and reasoning about reward design, credit assignment, and why feedback-driven systems become unstable.
What you'll do- Build and maintain the training, evaluation, and deployment loops at the core of the self-improvement system, with a strong emphasis on reproducibility and reliability.
- Design and implement reward and feedback signals; investigate and mitigate reward hacking, specification gaming, and distribution drift.
- Build evaluation harnesses and metrics before models - because a self-improving system is only as safe as its measurement of "better."
- Own data pipelines and automated data flywheels that feed the learning loop.
- Debug subtle model-quality regressions and stabilize training and feedback loops that go non-stationary.
- Collaborate with research and product to turn methods into robust, shippable systems.
What we're looking forMust-Haves:Strong MLE fundamentals (non-negotiable)- Excellent Python and clean, well-tested ML training code.
- Solid grasp of data pipelines, distributed / large-scale training, and experiment tracking.
- The instinct and skill to debug why a model silently got worse - not just why it crashed.
Hands-on experience with a feedback or learning loop (at least one).- Built or owned part of a feedback loop - a reward model, an evaluation harness, or the data pipeline for an RLHF/RLAIF or active-learning system.
- Ran a retraining or continual-learning pipeline where a model consumed its own predictions or production data (e.g. ranking, recommendations, fraud, spam).
- Fine-tuned LLMs with human or AI feedback, or built agentic evaluation harnesses.
Reinforcement learning foundations and curiosity.- Working knowledge of reward modeling, on-policy vs. off-policy tradeoffs, and credit assignment (does not need to be a research-level RL expert).
- Has seen - or can reason clearly about - feedback-system failure modes: reward hacking, specification gaming, feedback loops amplifying errors.
- Comfortable evaluating non-stationary systems (systems whose behavior and data distribution change over time).
Systems and evaluation instinct.- Builds the eval before the model; treats measurement as a first-class deliverable.
- Has shipped an ML system into production and kept it healthy over time.
Strongest signal Bonus, not required:The ideal candidate has built or shipped a full system that improved from its own outputs or feedback end to end. This is rare at this level, so treat it as a standout differentiator rather than a filter. Examples:
- RLHF / RLAIF pipelines
- Self-play systems
- Active-learning loops
- Automated data flywheels
- Agentic evaluation harnesses
Nice to Have:- PhD or MS in RL / ML paired with real production experience (either the science or the engineering half alone is fine if the other is strong).
- Experience at a lab or company doing RLHF, agents, or large-scale ML infrastructure.
- Familiarity with LLM fine-tuning, evaluation frameworks, or agent orchestration.