About the RoleThis is a high-ownership, high-impact role at an early-stage company. You'll work in ambiguous, fast-moving problem spaces alongside a tight-knit team where your contributions directly influence the trajectory of frontier AI development.
What You'll Do- Build systems for creating new environments, improving data quality, and translating real-world workflows into tasks and benchmarks.
- Build systems for creating, running, evaluating, and improving agent training environments.
- Design experiments to understand model behavior, agent failure modes, and data quality issues.
- Develop tools that help researchers, engineers, and data vendors create higher-quality tasks, trajectories, and feedback loops.
- Work across the full lifecycle of agent training data - from task design and environment setup to trajectory collection, evaluation, and validation.
- Partner with external vendors to identify bottlenecks and improve the quality and throughput of the data engine.
- Build metrics and analyses to assess whether tasks, environments, and evals are genuinely useful for training frontier agents.
What We're Looking ForRequired:- 2-4 years of relevant engineering experience.
- Proficiency in Python, Docker, and Linux environments.
- Experience with benchmarks and evals, including reasoning about task realism, rubric reliability, environment usability, and trajectory quality for RL training.
- Strong attention to detail - ability to spot subtle inconsistencies in data, model behavior, or task design.
- Track record of building tools, pipelines, or research infrastructure with minimal guidance.
- Early-stage startup experience; comfort working independently in fast-paced, ambiguous settings.
- Experience designing metrics and validation workflows.
- Strong quantitative or technical foundation, demonstrated through competitive programming, research, or independent project work.
- Ability to thrive in unstructured problem spaces and communicate clearly across time zones.
Nice to Have:- Background in reinforcement learning or AI alignment research.
- Experience working with large-scale data pipelines or vendor ecosystems.
- Publications or contributions to open-source ML/AI tooling.
Compensation & Benefits- Salary: $150,000 - $250,000 USD annually, depending on experience.
- Visa sponsorship is available.
- Equity participation in an early-stage, well-resourced AI company.
LocationThis is an
on-site role based in
San Francisco, CA. Candidates should be prepared to work in-person with the team.