Member of Technical Staff (MTS) - Multimodal Foundation Models

Deeproute.ai

$130K — $180K *
Transportation
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • MS or PhD in Computer Vision, Machine Learning, Robotics, Computer Science, or related fields
  • Thorough understanding of foundation models, self-supervised and multimodal learning, and large-scale pretraining
  • Hands-on experience with CLIP, DINO/DINOv2, MAE, contrastive learning, masked modeling, or scalable transformer architectures
  • Desirable experience in video foundation models, long-context modeling, retrieval systems, and efficient inference
  • Proven track record with publications in top-tier venues like CVPR, ICCV, ECCV, NeurIPS, ICLR, or ICML

Responsibilities

  • Develop scalable pretraining pipelines for multimodal driving data
  • Design and optimize training strategies for vision-language-action and video foundation models
  • Enhance training stability, data efficiency, and representation robustness
  • Conduct architecture-level research into Vision Transformers and multimodal systems
  • Explore and improve pretraining objectives and training paradigms
  • Analyze model performance through rigorous ablation studies and failure case analysis
  • Improve the efficiency and deployability of multimodal foundation models for large-scale production

Benefits

  • Collaborative environment with a strong emphasis on innovation
  • Opportunity to work on cutting-edge AI technologies for real-world applications
  • Engagement in research that bridges academic knowledge and practical development
  • Access to advanced computational resources and distributed training frameworks
  • Potential for a significant impact in the field of autonomous driving systems
Full Job Description
Focus

Multimodal Foundation Models • Representation Learning • Method Innovation

We are looking for strong technical builders and researchers who deeply understand foundation models and representation learning beyond simply applying existing frameworks.

Ideal candidates should have:
  • Strong experimental rigor
  • Solid systems and modeling intuition
  • Hands-on engineering ability
  • Interest in scalable multimodal AI systems for real-world autonomy

We value people who can bridge research and production, and who care about robustness, scalability, efficiency, and practical deployment in large-scale autonomous driving systems.

Responsibilities

1. Large-Scale Foundation Model Pretraining
  • Develop scalable pretraining pipelines for large-scale multimodal driving data
  • Design and optimize training strategies for:
      • Vision-language-action models
      • Video foundation models
      • Long-context temporal modeling
      • Multimodal representation alignment
  • Improve:
    • Training stability
    • Data efficiency
    • Scaling efficiency
    • Representation robustness
  • Work on distributed training systems and large-scale model optimization using frameworks such as:
    • PyTorch Distributed
    • DeepSpeed
    • Megatron-LM

2. Representation Learning & Method Innovation
  • Design and improve self-supervised and multimodal learning methods for real-world autonomous driving systems
  • Conduct architecture-level research on:
    • Vision Transformers (ViT)
    • Video / temporal architectures
    • Multimodal fusion and alignment
    • Embedding and retrieval systems
    • Long-context and memory-efficient architectures
  • Explore and improve:
    • Pretraining objectives
    • Loss functions
    • Training paradigms
    • Generalization and robustness
  • Analyze model behavior through:
    • Rigorous ablation studies
    • Failure case analysis
  • Representation probing and evaluation

3. Efficient Foundation Models & Scalable Deployment
  • Improve the efficiency, scalability, and deployability of large multimodal foundation models for real-world autonomous driving systems
  • Work on areas such as:
    • Model quantization
    • Knowledge distillation
    • Efficient attention mechanisms
    • Sparse architectures and Mixture-of-Experts (MoE)
    • Long-context and memory-efficient modeling
    • Inference acceleration and serving optimization
    • Training and inference system efficiency
  • Optimize model throughput, latency, memory usage, and deployment performance for large-scale production environments

Requirements
  1. MS or PhD in:
      • Computer Vision
      • Machine Learning
      • Robotics
      • Computer Science
      • Related fields
  2. Strong understanding of:
      • Foundation models
      • Self-supervised learning
      • Representation learning
      • Multimodal learning
      • Large-scale pretraining
  3. Hands-on experience with methods such as:
      • CLIP
      • DINO / DINOv2
      • MAE
      • Contrastive learning
      • Masked modeling
      • MoE or scalable transformer architectures
  4. Experience with one or more of the following is highly valued:
      • Video foundation models
      • Long-context modeling
      • Retrieval systems
      • Efficient inference
      • Distributed training
      • Model compression and deployment optimization
  5. Strong publication record in top-tier venues is preferred:
      • CVPR
      • ICCV
      • ECCV
      • NeurIPS
      • ICLR
      • ICML

Similar Jobs

More Jobs at Deeproute.ai

More Transportation Jobs

Find similar Member of Technical Staff (MTS) - Multimodal Foundation Models jobs: