About The RoleInvisible is building a reinforcement learning capability inside its research organization, focused on evaluation methodology, benchmarks, and RL environments for frontier labs and enterprise clients. As a Research Engineer, you'll work on how frontier models are measured, meaning the evaluations that labs and enterprises actually make decisions on, with direct influence over the methodology and not just its implementation.
You'll take a research question (how do we measure whether a model can do this work, and how do we make that measurement reproducible?) and turn it into a scoring framework, an evaluation architecture, or an environment design, then ship the production system that runs it. This role doesn't hand specifications to someone else to build, and it doesn't only build what others have specified. It's a fit for someone who wants to design the approach and write the code that proves it out.
Depending on level, the role reports to the Lead Research Engineer or, at Principal, directly to the VP of Research.
What You'll Do- Design benchmarks and RL environments that measure real model capability for frontier labs and enterprise clients
- Originate evaluation methodology and translate it into scoring frameworks, rubrics, and evaluation architectures
- Write and ship the production code that implements your designs
- Build and maintain the data pipelines that feed evaluation runs
- Run the analysis that establishes whether a result holds, and make evaluations reproducible
- Partner with Research Scientists on methodology review, with Solutions Architects on client requirements, and with ML software engineers who build and maintain the underlying platform
What We Need- Production-quality code written daily; this is a hard requirement
- Python & ML Stack: Fluency in Python and comfort across the modern ML stack
- Evaluation & Infrastructure: Real experience building evaluation systems, RL environments, or training and inference infrastructure
- Agentic Systems: Hands-on experience with modern agentic flows
- RL & Frontier Evaluation: Familiarity with reinforcement learning methods and how frontier models are evaluated; we weigh this more heavily than years of experience
- A track record of turning ambiguous research questions into working systems, and publishing or shipping the result
You can find more information about our geographic pay tiers
here. During the interview process, your Invisible Talent Acquisition Partner will confirm which tier applies to your location. For candidates outside the U.S., compensation is adjusted to reflect local market conditions and cost of living.
*Bonuses and equity are included in all full-time offers. Final compensation is determined by a combination of factors, including location, job-related experience, skills, knowledge, internal pay equity, and overall market conditions. Because of this, every offer is unique. Additional details on total compensation and benefits will be discussed during the hiring process.