Hands-on experience with AI environments or reinforcement learning infrastructure.
Strong software engineering fundamentals.
Proven track record of building and deploying technical infrastructure.
Understanding of evaluation methods, reward structures, and agent trajectories.
Proficiency in multiple programming languages and technology stacks.
Responsibilities
Design evaluation environments for long-horizon enterprise workflows.
Define essential tasks, states, tools, and reward signals for agent evaluation.
Build high-fidelity representations of complex enterprise software environments.
Develop infrastructure for agent rollouts and trajectory inspection.
Measure correctness and efficiency of multi-step agent behaviors.
Investigate evaluation failures and issues related to reward quality.
Create production-quality systems instead of research prototypes.
Benefits
Opportunity to work on cutting-edge AI technologies.
Collaborative work environment with a focus on engineering principles.
Flexible arrangements considered for exceptional candidates.
Exposure to complex enterprise software and real-world applications.
Full Job Description
Research Engineer - RL Infrastructure & Agent Environments
San Francisco, California | Primarily On-site
We are seeking an Research Engineer - RL Infrastructure & Agent Environments to build the environments, evaluation systems, and supporting infrastructure used to train and assess long-horizon enterprise AI agents.
The Opportunity
You will work on the engineering and research problems behind realistic agent environments, post-training systems, and reliable evaluation of complex multi-step workflows.
Key Responsibilities
Design evaluation environments for long-horizon enterprise agent workflows.
Define tasks, state, tools, graders, and reward signals used to evaluate and improve agents.
Build high-fidelity representations of complex enterprise software environments.
Develop infrastructure for rollouts, orchestration, trajectory inspection, and grader pipelines.
Measure both correctness and efficiency across multi-step agent behavior.
Investigate evaluation failures, reward-quality issues, and agent behavior.
Build production-quality systems rather than notebook-only research prototypes.
Required Qualifications
Hands-on experience with AI environments, evaluations, reinforcement learning infrastructure, or related agent-training systems.
Strong software engineering fundamentals.
Demonstrated ability to build and ship technical infrastructure.
Understanding of evaluation methodology, reward design, graders, and agent trajectories.
Ability to work across languages and technology stacks based on system requirements.
Candidate Profile
A PhD is not required. Strong engineering and shipped environment or evaluation systems are more important than academic credentials or publication history.
Seniority
The opportunity is open to exceptional new graduates, early-career engineers, and experienced senior candidates. Selection is based primarily on engineering strength and relevant technical work.
Work Arrangement
The role is anchored in San Francisco with a strong preference for in-person collaboration. Limited flexibility may be considered case by case for exceptional candidates.