The Impact You'll MakeOur research team is expanding to keep pace with a wave of frontier-facing work: internal research streams, client engagements that require real ML depth, and emerging opportunities at the cutting edge of the field. As a Research Engineer, you'll take a research direction and run with it - finding the right papers, benchmarks, and prior work, reimplementing what's relevant, and building out the process to reproduce and improve on it internally.
You'll own initiatives end to end: partnering with strategic project and technical leads to scope the work, building MVPs to validate ideas (including through human annotation and agents), and turning that work into something concrete - a customer dataset, a pilot, an internal dataset that becomes a paper or blog post, or a joint publication with a partner. You won't be handed a fully specified task list; you'll be given a direction and the autonomy to turn it into a research plan.
This is a full-time, hybrid position based in San Francisco.
What You'll Do- Take a research direction and independently identify supporting resources - papers, benchmarks, blog posts - then implement or reimplement the relevant methods.
- Build and own the process to reproduce prior work internally and identify ways to improve on it.
- Own projects (for example, an RL/agentic environment build for a partner or a novel multimodal benchmark) end to end, including scoping, MVP implementation, and validation.
- Partner with strategic project leads and technical leads to translate ambiguous requirements into a concrete, testable research plan.
- Validate ideas through hands-on implementation, including annotating, evaluating, or sourcing data.
- Turn research directions into tangible outputs - a paid customer dataset, a customer pilot, an internal dataset, or a paper/blog post for publication or conference presentation.
- Bring an ML perspective to new opportunities - assessing technical feasibility of incoming requests and helping shape proposals where research depth is needed.
What You'll Bring- MS or PhD in ML, CS, or a related quantitative field - or equivalent demonstrated research experience (publications, significant open-source research work, industry research).
- Real ML depth: you understand how models are trained and evaluated, not just how to call an API. You can read a paper, judge whether its claims hold, and reimplement the method.
- Hands-on experience with at least one of: RL/agentic systems, AI/ML evaluation and benchmarking, or multimodal ML.
- Strong Python and the engineering ability to build and ship your own experiments - eval harnesses, environments, infrastructure - without relying on a platform team.
- High autonomy: you can turn an ambiguous direction into a concrete research plan and notice when something's off before being told.
- Clear technical writing
Nice To Have- Publication track record (first-author preferred).
- Experience with agent or multimodal benchmarks (OSWorld, MMMU, WebArena, SWE-bench, or similar) or building RL environments/gyms.
- Familiarity with reward modeling, reward hacking, or verifier/judge reliability.
- Familiarity with synthetic data generation or human-in-the-loop (HITL) workflows.
- Experience with cloud infrastructure and containerized environments.
- A deep RL background specifically.
$180,000 - $280,000 a year
In addition to the annual base salary, employees are eligible for an annual bonus paid out quarterly.
Only shortlisted candidates will be contacted for an interview!