What youll doWere building a replayable environment engine over real enterprise history.
The system reconstructs a companys context as it existed at any past time, then exposes that state through the same tools an agent would use in production. This lets us place new policies and agent configurations inside real historical environments, observe how they reason and act, and grade their performance against real outcomes.
Youll own the research and infrastructure required to turn this into a scalable model-improvement system. The core problems include:
- Building an environment factory that converts recorded enterprise data and task definitions into runnable environments
- Designing graders that turn ambiguous business objectives into verifiable rewards
- Developing methods for mining useful tasks, trajectories, and evaluation cases from historical workflows
- Creating eval sets that are representative, reproducible, and resistant to overfitting
- Finding the right combinations of models, tools, context, and policies to maximize performance while reducing inference cost
- Training and evaluating agents that operate over long horizons, incomplete information, and large tool spaces
- Building replay and observability systems that make agent behavior explainable and measurable
- Scaling from individual environments to thousands of concurrent training and evaluation runs
These problems are wide open. Youll have significant ownership over both the research direction and the production systems that make it real.
Youll work directly with the CTO, deploy into real enterprise workflows, and see your research tested against consequential problems and observable outcomes.
Who you are- You have 4+ years of experience building production software or machine-learning systems, including at least 2 years working on reinforcement-learning environments, LLM post-training, evaluation infrastructure, agent harnesses, or closely related systems
- You understand how environment design, reward design, context, tooling, and policy behavior interact
- Youre comfortable turning fuzzy business objectives into tasks and signals that can be evaluated reliably
- You can diagnose whether a models limitations come from the model itself, its context, its tools, its harness, or its training
- You can move between research questions and production implementation without treating them as separate jobs
- You write strong software and can build systems that process large, messy datasets at scale
- You care about reproducibility, observability, and understanding why a model behaves the way it does
- Youre looking to do the best work of your life and build something youll be proud of for decades
We care much more about what youve built and how you think than credentials or conventional career paths.
Benefits- Significant equity and ownership
- Equinox membership
- Free meals, coffee, and snacks
- Health insurance
- Unlimited PTO