OverviewYou will design and run training experiments that isolate the impact of our datasets on model behavior. Through controlled SFT and RL - base post-training experiments, you will measure how different data sources, structures, and selection strategies affect capability, generalization, and alignment. You will own the full experiment loop: formulate the hypothesis, build the pipeline, run the model, analyze the results, identify confounders, and determine what should happen next. Working with partner labs and internal teams, you will turn our datasets into clear, defensible evidence: this data -> this improvement -> under these conditions. This is empirical, high-leverage work for someone who likes building quickly and extracting signal from noisy results.
ResponsibilitiesDesign and build our post-training infrastructure for SFT, RL and evaluation workflows.
Build reliable pipelines for data preparation, dataset versioning, sampling, training, checkpoint management, and evaluation.
Develop experiment orchestration and tracking systems that make runs reproducible, comparable, and easy to debug.
Create reusable abstractions that allow researchers to launch experiments quickly across datasets, models, and training recipes.
Integrate our systems with partner-lab training stacks, model APIs, compute environments, and evaluation infrastructure.
Required QualificationsAt least 2 years of professional experience in machine learning engineering, research engineering, ML infrastructure, or a closely related field; 2-4+ years preferred.
Strong Python and software-engineering skills.
Hands on experience with PyTorch, JAX, Ray, and Slurm.
Experience building production-quality ML training, evaluation, or data infrastructure.
HAnds-on experience with LLM fine-tuning, post-training, and evaluation.
Ability to build reliable, reproducible systems for launching and comparing ML experiments.
Strong debugging skills across distributed systems, data pipelines, training infrastructure, and model behavior.
Understanding of experimental design and the ability to extract actionable conclusions from noisy results.
Ability to move quickly between infrastructure engineering and hands-on experimentation.
A bias toward building, testing, and shipping.
Company Benefits (For Eligible Employees):- Health Insurance
Medical, Vision, Dental - 401(k)
With Employer Match - Daily Meals
Daily UberEats Stipend - Wellness Stipend
Monthly - Covers Equinox Membership - Commute Covered