The RoleIn RL, quality is not a polish step-it is the training signal itself.
When task designs are ambiguous, sandboxes lack proper isolation, or reward functions contain loopholes, agents do what optimizers always do: exploit the grader. A broken environment doesn't just burn cluster hours; it actively poisons downstream model weights.
As Head of Quality, you will own the standard for what constitutes real learning signal across every task, verifier, and environment Plato ships. You'll be the adversarial mind finding the loopholes before the model does, the engineer automating the test suites, and the leader directing the team running verification.
What You'll Do- Gate Final Delivery: Own the final sign-off before environments and datasets ship to frontier labs, auditing trajectories, tasks, and reward dynamics for correctness, feasibility, and signal density.
- Harden Verifiers & Tasks: Build automated red-teaming suites to stress-test task feasibility and verifier integrity, aggressively eliminating reward hacking, grader tampering, and impossible task traps.
- Automate Verification Infra: Architect judge models, sandbox replay harnesses, and rollout forensics to catch synthetic drift, out-of-distribution behaviors, and leaky states upstream.
- Build & Lead the Verification Team: Hire and direct a high-agency team of QA engineers, domain specialists, and technical reviewers, blending automated agentic checks with deep human-in-the-loop review.
- Close the Loop with Research: Translate downstream model failure modes into concrete generator constraints so defects are prevented at the generation stage rather than caught in review.
Required Qualifications- Technical Depth in ML/RL or Systems: 3+ years of experience in software engineering, ML engineering, or research systems, with proficiency in modern languages (Python, etc.).
- Adversarial Systems Instinct: You understand how optimizers exploit edges. You naturally think about how an agent could game a reward function, break a sandbox, or fake completion.
- Trajectory & Code Forensics: Proven ability to dive into raw rollouts, agent reasoning traces, and verification code to spot subtle ungrounded assumptions or hallucinated logic.
- High-Agency Leadership: Experience building or leading a technical evaluation, QA, or data verification function from zero to one.
- Zero-Compromise Bar: Comfort holding the line on delivery under intense customer pressure. You understand that shipping contaminated data is far worse than shipping late.
Preferred Qualifications- Experience with RL training dynamics, automated LLM evals, agent sandboxing, or synthetic trajectory generation.
- Background designing adversarial test suites, code execution verifiers, or formal verification systems.
- Experience handling client-facing technical evaluations and failure postmortems with frontier AI research teams.Introduction
Plato is an applied research lab building the foundational infrastructure to train specialized AI agents.
We turn real-world data streams into high-fidelity simulated environments that generate the training signal needed to make capable models. Our work supports frontier labs, hyperscalers, and enterprises building AI systems for complex, high-stakes work.
Today, only a handful of players can train models for capable work. Compute and algorithms are rapidly commoditizing, but reinforcement learning data remains the bottleneck. Plato is changing that by automatically scaling training environments from proprietary real-world data.