1-7 years of experience in production software or machine-learning systems
Familiarity with reinforcement-learning environments or related systems
Understanding of environment and reward design interactions
Ability to translate ambiguous business objectives into evaluable tasks
Skill in diagnosing model limitations across various components
Proficiency in processing large datasets at scale
Commitment to reproducibility and model behavior understanding
Responsibilities
Build an environment factory to convert enterprise data into runnable environments
Design graders for translating business objectives into measurable rewards
Develop methods for mining tasks from historical workflows
Create evaluation sets that are reproducible and prevent overfitting
Optimize models, tools, and policies for performance and cost efficiency
Train and evaluate agents operating over long horizons and incomplete information
Build systems for replay and observability of agent behavior
Benefits
Significant equity and ownership
Equinox membership
Free meals, coffee, and snacks
Health insurance
Unlimited PTO
Full Job Description
What you9ll do
At the center of Ambral Labs is a replayable environment engine for the enterpirse.
The system reconstructs a company9s world as it existed at a particular moment in the past then exposes that state through the same tools an agent would use in production. This lets us place new policies and agent configurations inside real historical environments, observe how they reason and act, and grade their performance against real outcomes.
You9ll work across research, infrastructure, and production systems including:
Building an environment factory that converts recorded enterprise data and task definitions into runnable environments
Designing graders that turn ambiguous business objectives into verifiable rewards
Developing methods for mining useful tasks, trajectories, and evaluation cases from historical workflows
Creating eval sets that are representative, reproducible, and resistant to overfitting
Finding the right combinations of models, tools, context, and policies to maximize performance while reducing inference cost
Training and evaluating agents that operate over long horizons, incomplete information, and large tool spaces
Building replay and observability systems that make agent behavior explainable and measurable
Scaling from individual environments to thousands of concurrent training and evaluation runs
These problems are wide open. You9ll have significant ownership over both the research direction and the production systems that make it real.
You9ll work directly with the CTO, deploy into real enterprise workflows, and see your research tested against consequential problems and observable outcomes.
Who you are
You have 1-7 years of experience building production software or machine-learning systems (we9re hiring at multiple levels for this role).
Bonus points for working on reinforcement-learning environments, LLM post-training, evaluation infrastructure, agent harnesses, or closely related systems
You understand how environment design, reward design, context, tooling, and policy behavior interact
You9re comfortable turning fuzzy business objectives into tasks and signals that can be evaluated reliably
You can diagnose whether a model9s limitations come from the model itself, its context, its tools, its harness, or its training
You can move between research questions and production implementation without treating them as separate jobs
You write strong software and can build systems that process large, messy datasets at scale
You care about reproducibility, observability, and understanding why a model behaves the way it does
You9re looking to do the best work of your life and build something you9ll be proud of for decades
We care much more about what you9ve built and how you think than credentials or conventional career paths.