Pre-training Research Engineer

Sciforium

$150K — $180K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in machine learning research or engineering, specifically with large language or multimodal models.
  • Strong software engineering skills for writing effective training code.
  • Solid understanding of deep learning fundamentals and modern pre-training techniques.
  • Experience in implementing and evaluating research ideas with clear baselines and analysis methods.
  • Hands-on background in GPU and distributed training environments.
  • MS degree in Computer Science, Machine Learning, AI, Mathematics, or related field.

Responsibilities

  • Train large byte-native and multimodal foundation models using diverse data.
  • Implement and assess new model architectures and training objectives.
  • Develop effective pre-training recipes and conduct scaling experiments.
  • Perform ablation studies and analyze model quality and training dynamics.
  • Collaborate with data engineers to enhance training efficiency and scalability.

Benefits

  • Medical, dental, and vision insurance.
  • 401k plan.
  • Daily lunch, snacks, and beverages provided.
  • Flexible time off.
  • Competitive salary and equity options.
Full Job Description
About the role

As a Pre-training Research Engineer, you'll focus on model implementation, pertaining and scaling, and improving the quality of our byte-native and multimodal foundation models. You'll build and iterate quickly on research ideas, contribute production-grade training code and infrastructure, and help deliver high-quality base models that can serve real-world use cases at scale.
Key Responsibilities
Pre-training & Scaling
  • Train large byte-native and multimodal foundation models across massive, heterogeneous corpora.
  • Implement and evaluate new model architectures, training objectives, and optimization methods.
  • Develop stable pre-training recipes and run scaling experiments for novel architectures.
  • Conduct ablations and analyze training dynamics, model behavior, and base-model quality.
  • Work with data and distributed training engineers to improve training efficiency, reliability, and scalability.
Must-Haves
  • 5+ years of experience in machine learning research or engineering, with a proven track record of developing and pre-training large language or multimodal foundation models.
  • Software Engineering: Strong general software engineering skills, with the ability to write robust and performant training code.
  • ML Foundations: Solid understanding of deep learning fundamentals and modern pre-training methods and literature.
  • Research and Experimentation: Ability to quickly implement research ideas and evaluate them using clear baselines, ablations, metrics, and analysis.
  • GPU and Distributed Training: Hands-on experience running training workloads in GPU-based environments, with familiarity with distributed training.
  • Education: MS in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.
Nice-to-Haves
  • PhD in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.
  • JAX Ecosystem: Extensive experience with the JAX, Flax, and XLA stack.
  • Large-Scale Distributed Training: Experience with multi-node pre-training using systems such as FSDP, ZeRO, or Megatron.
  • Training Recipes and Scaling: Experience developing training recipes, ablations, or scaling experiments.
  • Monitoring and Reproducibility: Experience owning end-to-end training and evaluation pipelines with monitoring and reproducibility.
Education
  • MS or PhD in Computer Science, Machine Learning, Artificial Intelligence, Mathematics, or a related field.


Benefits include
  • Medical, dental, and vision insurance
  • 401k plan
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary and equity


Similar Jobs

More Jobs at Sciforium

More Information Technology Jobs

Find similar Pre-training Research Engineer jobs: