Staff ML Engineer, VLA Training

Nidus Technologies

$150K — $180K *
Consumer Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in machine learning, robotics, computer vision, or similar fields.
  • Proficient in training models with PyTorch, JAX, or similar frameworks.
  • Experience with transformer-based, multimodal, and generative models.
  • Skilled in building distributed training pipelines on GPU or multi-node clusters.
  • Knowledge of large-scale datasets, data loading, and reproducible ML workflows.
  • Strong understanding of optimization, model architecture, and experimental design.
  • Clear communication skills and ability to collaborate across various teams.

Responsibilities

  • Train and enhance vision-language-action models for robotic tasks.
  • Develop model architectures and training recipes using various machine learning techniques.
  • Build scalable pipelines for model pretraining and evaluation.
  • Train models across multi-node GPU clusters while optimizing performance and cost.
  • Create data mixtures and sampling strategies for robot datasets.
  • Design systems for managing video, language, and sensor data.
  • Run experiments on physical robots to inform improvements in models and infrastructure.

Benefits

  • Work at the forefront of robotics and machine learning.
  • Collaborate in a multidisciplinary team of engineers and researchers.
  • Opportunity to mentor and establish engineering practices in the ML domain.
  • Engage in hands-on projects that drive the advancement of intelligent robots.
Full Job Description
About the role

We are building intelligent robots that can perceive the world, understand instructions, and perform useful physical tasks in dynamic, real-world environments.

We are looking for a Staff Machine Learning Engineer to develop and scale the vision-language-action models that power our robots. You will own major parts of the model development lifecycle-from dataset construction and training recipes to distributed training infrastructure, evaluation, and on-robot validation.

This is a hands-on role at the intersection of machine learning research, large-scale systems, and robotics. You will work closely with researchers, robotics engineers, data teams, and operators to turn new modeling ideas into reliable robot capabilities. The right person is equally comfortable investigating why a policy failed on a robot, designing the next training experiment, and improving the infrastructure required to run that experiment efficiently at scale.

What you'll do
  • Train and improve vision-language-action models for robotic manipulation and other embodied tasks.
  • Develop model architectures and training recipes spanning imitation learning, behavior cloning, transformer- and diffusion-based policies, multimodal pretraining, and post-training.
  • Build scalable pipelines for pretraining, fine-tuning, evaluation, checkpointing, and model release.
  • Train models across multi-node GPU clusters while improving utilization, throughput, stability, and cost efficiency.
  • Design data mixtures, sampling strategies, augmentations, and curriculum approaches for large, heterogeneous robot datasets.
  • Develop data loaders and preprocessing systems for synchronized video, language, robot state, actions, and other sensor modalities.
  • Create reproducible experimentation systems, including configuration management, dataset and model versioning, experiment tracking, and automated regression testing.
  • Define offline and on-robot metrics that measure task success, generalization, robustness, latency, safety, and failure modes.
  • Build tools for inspecting trajectories, visualizing model behavior, comparing experiments, and diagnosing data or training failures.
  • Run structured experiments on physical robots and use the results to guide model, data, and infrastructure improvements.
  • Partner with data collection teams to identify coverage gaps, improve demonstration quality, and prioritize collection based on model performance.
  • Translate promising research into maintainable, production-quality systems that can support repeated deployment across robots and environments.
  • Make technical decisions across model architecture, data, compute, and evaluation, and communicate the associated tradeoffs clearly.
  • Mentor other engineers and establish strong engineering practices for the ML codebase.
What we're looking for
  • Five or more years of professional experience in machine learning, robotics, computer vision, or a closely related field, or equivalent demonstrated impact.
  • Strong experience training modern deep learning models using PyTorch, JAX, or a comparable framework.
  • Experience with transformer-based models, multimodal models, generative models, or learned control policies.
  • Experience building and operating distributed training pipelines on multi-GPU or multi-node accelerator clusters.
  • Experience with large-scale datasets, high-throughput data loading, experiment tracking, and reproducible ML workflows.
  • Strong understanding of optimization, model architecture, data quality, evaluation methodology, and experimental design.
  • Ability to debug failures across the full stack, from input data and training dynamics to inference and robot behavior.
  • Experience independently owning technically ambiguous projects and delivering working systems.
  • Clear communication skills and an ability to collaborate across research, infrastructure, robotics, hardware, and operations teams.
  • A bachelor's degree in computer science, machine learning, robotics, electrical engineering, or a related field-or equivalent practical experience.
Nice to have
  • Experience developing vision-language-action models, vision-language models, or robotics foundation models.
  • Experience with imitation learning, reinforcement learning, behavior cloning, diffusion policies, action tokenization, or action-conditioned world models.
  • Experience training policies using data from multiple robot embodiments, task domains, or sensor configurations.
  • Familiarity with large-scale pretraining, post-training, parameter-efficient fine-tuning, distillation, or model adaptation.
  • Experience optimizing distributed training using techniques such as FSDP, tensor or pipeline parallelism, mixed precision, gradient checkpointing, or sharded data loading.
  • Familiarity with robotics middleware and simulation environments such as ROS/ROS 2, MuJoCo, Isaac Sim, or Gazebo.
  • Experience deploying and evaluating learned policies on physical robotic systems.
  • Publications or meaningful open-source contributions in machine learning, robotics, computer vision, or related areas.


Similar Jobs

More Jobs at Nidus Technologies

  • Senior Recruiter
    $110K — $130K *
    New York, NY 10025 (New York County)
    Technical Services
    In-Person
  • Business Operations Lead
    $110K — $130K *
    New York, NY 10025 (New York County)
    Business Services
    In-Person
  • Backend Engineer
    $135K — $160K *
    New York, NY 10025 (New York County)
    Manufacturing & Automotive
    In-Person

More Consumer Technology Jobs

Find similar Staff ML Engineer, VLA Training jobs: