Research Engineer, Post-Training DataLocation: San Fran
Employment Type: Full-time
Focus: Post-Training, Reinforcement Learning, Data Generation, Research Environments, Frontier AI
Profile: Researcher-engineer with strong ML and software engineering depth
About the RoleOur client is hiring a
Research Engineer, Post-Training Data to own the full lifecycle of post-training model and data work.
This role blends AI research, ML engineering, software engineering, and data production into one function. The ideal candidate is not just a researcher and not just an engineer - they are someone who can understand a domain deeply, identify what makes a task realistic and economically valuable, build the environment or data-generation system, and improve the model through high-quality post-training data.
The company is indexing heavily on research and ML horsepower, strong software ability, high slope, and data taste.
What You'll Do- Own post-training workflows end to end, from infrastructure provisioning through data generation and curation
- Build and curate high-quality post-training data from model and environment generation processes
- Produce reinforcement learning environments and long-horizon tasks across complex domains
- Work on research-loop, science, chip, physics, and other technically deep task environments
- Build systems that scale data generation superlinearly through self-improving developer processes
- Improve data generation through better tools, workflows, automation, evaluation, and model feedback loops
- Exercise and develop strong data taste around what problems matter in a given domain
- Identify realistic, economically valuable tasks and environments that frontier AI labs will actually want
- Blend research and data production into a single technical function
- Debug ML pipelines, understand unfamiliar code quickly, and solve open-ended technical problems
- Help define the standards for what high-quality post-training data should look like across domains
What We're Looking For- Genuine spike in ML, AI research, or a technical research domain
- Experience or strong interest in post-training, reinforcement learning, evaluations, interpretability, or related areas
- Strong software engineering ability and comfort building production-quality tools or research systems
- Ability to post-train models and build or curate high-quality post-training data
- Strong judgment around data quality, task design, and what makes a problem valuable to AI labs
- Fast problem-solving ability and strong code comprehension
- Ability to work across unfamiliar domains and ramp quickly
- High-slope learning profile with strong technical curiosity
- Comfort operating in ambiguous, research-heavy environments
- Evidence of deep commitment, strong output, and public artifacts or meaningful technical work
Ideal BackgroundOur client is especially interested in candidates who combine:
- Domain or research expertise in a field such as neuroscience, physics, chemistry, systems, chip design, science, or another technical discipline
- Strong ML or software engineering core
- Experience with PyTorch, ML pipelines, model debugging, or research tooling
- Experience building environments, evaluations, or data-generation systems
- Exposure to post-training, RL, long-horizon tasks, or frontier model workflows
Pedigree can be a helpful signal, but it is not the primary filter. The team cares more about slope, research horsepower, problem-solving speed, and the ability to build.
Bonus Experience- Experience at frontier AI labs, AI infrastructure companies, or high-talent technical teams
- Experience with post-training data, RL environments, evals, model behavior, or interpretability
- Experience building data products, research tooling, or developer workflows for AI teams
- Experience in domains like chip design, physics, chemistry, biology, systems, or scientific computing
- Public artifacts, research projects, open-source work, technical writing, or demos that show exceptional ability
- Experience creating tasks or environments that require long-horizon reasoning
- Experience scaling data generation through automation rather than manual labor
Who Will Thrive Here- A researcher-engineer who treats data as a research problem
- A high-slope polymath with a real technical spike
- Someone who can figure out unfamiliar domains quickly
- A builder who understands that the best data comes from strong research judgment
- Someone obsessed with creating realistic, economically valuable AI training and evaluation environments
- A deeply committed, results-driven technical operator
- Someone excited by frontier lab customers and the thesis that better systems can scale data generation superlinearly
Why This Opportunity- Join an early team serving fast-growing demand from frontier AI customers
- Work on the core bottleneck behind better post-training outcomes: high-quality data
- Build RL environments and long-horizon task systems across technically deep domains
- Shape how research and data production come together as one function
- Work on systems designed to scale data generation through process, tooling, and model feedback - not simply more people
- Build at the frontier of post-training, evaluations, RL environments, and data quality
- Step into a role where research taste, engineering speed, and domain curiosity all matter
Ideal Candidate ProfileThe ideal candidate is a research-minded engineer with a real spike in ML, AI research, or a technical domain.
They can read and understand code quickly, debug ML pipelines, build tools, reason about task quality, and identify what makes a data environment valuable to a frontier AI lab. They are not looking for a narrow research role or a pure software role - they want to build the systems and data that make models better.
This person is high-slope, obsessive, technically broad, and excited to help define what great post-training data looks like.