About the roleAs a Pre-training Research Engineer, you'll focus on model implementation, pertaining and scaling, and improving the quality of our byte-native and multimodal foundation models. You'll build and iterate quickly on research ideas, contribute production-grade training code and infrastructure, and help deliver high-quality base models that can serve real-world use cases at scale.
Key ResponsibilitiesPre-training & Scaling- Train large byte-native and multimodal foundation models across massive, heterogeneous corpora.
- Implement and evaluate new model architectures, training objectives, and optimization methods.
- Develop stable pre-training recipes and run scaling experiments for novel architectures.
- Conduct ablations and analyze training dynamics, model behavior, and base-model quality.
- Work with data and distributed training engineers to improve training efficiency, reliability, and scalability.
Must-Haves- 5+ years of experience in machine learning research or engineering, with a proven track record of developing and pre-training large language or multimodal foundation models.
- Software Engineering: Strong general software engineering skills, with the ability to write robust and performant training code.
- ML Foundations: Solid understanding of deep learning fundamentals and modern pre-training methods and literature.
- Research and Experimentation: Ability to quickly implement research ideas and evaluate them using clear baselines, ablations, metrics, and analysis.
- GPU and Distributed Training: Hands-on experience running training workloads in GPU-based environments, with familiarity with distributed training.
- Education: MS in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.
Nice-to-Haves- PhD in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.
- JAX Ecosystem: Extensive experience with the JAX, Flax, and XLA stack.
- Large-Scale Distributed Training: Experience with multi-node pre-training using systems such as FSDP, ZeRO, or Megatron.
- Training Recipes and Scaling: Experience developing training recipes, ablations, or scaling experiments.
- Monitoring and Reproducibility: Experience owning end-to-end training and evaluation pipelines with monitoring and reproducibility.
Education- MS or PhD in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.
Benefits include- Medical, dental, and vision insurance
- 401k plan
- Daily lunch, snacks, and beverages
- Flexible time off
- Competitive salary and equity