RL Environments Engineer

BespokeLabs.AI, Inc

$250K — $300K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years in building and scaling agentic coding environments
  • Strong software engineering skills with experience across multiple languages
  • Demonstrated ability to automate and optimize coding tasks
  • In-depth knowledge of production software challenges and solutions
  • Ability to identify and mitigate potential cheating in grading systems
  • Proven track record of ship volume and effective coding environments
  • Strong sense of ownership and capability to work independently

Responsibilities

  • Build comprehensive coding environments based on complex real-world codebases
  • Select valuable coding environments that expose weaknesses in frontier models
  • Design tasks through the entire lifecycle, ensuring rigor and fairness
  • Develop robust grading systems that can withstand attempts to cheat
  • Enhance team efficiency by creating internal tooling solutions
  • Oversee coding agents to validate environments and detect failures
  • Iterate on coding tasks to improve clarity and effectiveness

Benefits

  • Health, dental, and vision coverage
  • 401(k) plan
  • Daily onsite lunch provided
  • Visa sponsorship and relocation support available
  • Opportunity to influence industry standards in agent training and evaluation
Full Job Description
About the Role

You will own coding environments end to end. You choose what world to build, design the tasks inside it, build the grading, run frontier models against it, and harden it until the only way to pass is to actually do the work.

This is a delivery role. We care most about whether you have done this before and can point to what came out of it. If you have built agentic coding environments or tasks at real volume and can tell us how many and how hard they were, we want to talk.

What You'll Do

Build high-fidelity coding worlds around real codebases, with the conventions, dependencies, tooling, and accumulated mess that real software has.

Choose which environments are worth building. A strong coding environment hits several marks:
  • Targets work where frontier models measurably struggle
  • Exercises real engineering, meaning navigation, diagnosis, sequencing, and design, and not just writing a function
  • Rests on a codebase with enough history and structure that shortcuts do not survive
  • Has a clear pass condition that a reviewer would agree with
  • Produces many varied tasks from a single world rather than one

Design tasks across the full lifecycle. Prompt, environment, grader, running frontier models, failure analysis, and iteration, until each task is rigorous, fair, and hard to game.

Build grading and sandboxed execution that is deterministic and cannot be gamed. Assume the model will try to pass without doing the work, and close the door before it finds it.

Remove whatever is slowing the team down. Build the internal tooling that makes everyone around you faster.

Direct frontier coding agents heavily to build and validate environments, judging their output and catching the quiet failures they produce.

What We're Looking For

A record of shipped volume. You have built agentic coding tasks or environments and can show us how many you personally drove and what they cost to produce.

Experience scaling that output through automation rather than through more people doing more manual work.

Strong software engineering fundamentals and fluency in several languages that holds up in production code.

Real experience with production software. Large codebases, build systems, testing, deployment, on-call, and root cause analysis. You know what real engineering work feels like because you have done it.

An adversarial mindset. You look at a grader and ask how a model would cheat it, and then you fix that.

A clear sense of what frontier coding agents can and cannot do, and where they cut corners.

Ownership. You build, debug, and ship without much supervision.

You May Be a Good Fit If You Also
  • Have worked on RL training systems, post-training, verifiers, or tool-use harnesses
  • Come from developer tooling, CI/CD, sandboxes, or code execution infrastructure
  • Have built large-scale automated test generation, fuzzing harnesses, or benchmark suites, which is close cousin work even if it was never called an RL environment
  • Have contributed to a public agentic benchmark such as Terminal-Bench
  • Have open-source work that other people depend on


What We Offer
  • Location: Mountain View, CA (Onsite)
  • Base Salary: $250,000 - $300,000 USD / year
  • Additional Comp: 25% performance-based bonus + equity

Benefits & Perks
  • Health, dental, and vision coverage
  • 401(k)
  • Daily onsite lunch provided
  • Visa sponsorship and relocation support available
  • Direct impact on how the industry trains and evaluates agents

We value different backgrounds and paths into this work. If this role excites you but you do not check every box, apply anyway.

Similar Jobs

More Enterprise Technology Jobs

Find similar RL Environments Engineer jobs: