Research Scientist, Multi-Modal Human Understanding

Meta

$145K — $175K *
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science or a related technical field, or equivalent experience
  • 2+ years of experience in multi-modal AI research, particularly with Vision-Language Models or video understanding
  • 2+ years of hands-on experience implementing large-scale neural networks, preferably with PyTorch
  • Proficient in designing experiments to assess multi-modal model performance and conducting quantitative analysis
  • Proficient in Python for developing production or research-quality code for multi-modal AI applications

Responsibilities

  • Design and implement innovative multi-modal architectures combining vision, language, and temporal data
  • Develop and train Vision-Language Models for various tasks including visual question answering and human-centric understanding
  • Build video foundation models for temporal reasoning and long-form video synthesis
  • Research generative synthesis techniques for producing human-centric video and multi-modal content
  • Conduct experiments to evaluate model performance and improve accuracy and generalization
  • Contribute to the complete research lifecycle from problem formulation to evaluation

Benefits

  • Access to cutting-edge AI research and multi-modal technologies
  • Collaboration with leading experts in the AI field
  • Opportunities for publishing in prestigious conferences in AI and computer vision
  • Support for continuous learning and professional development
  • Comprehensive health and wellness programs
  • Flexible work environment promoting work-life balance
Full Job Description
Meta is seeking a Research Scientist to advance multi-modal AI technologies for human understanding and synthesis. In this role, you will develop Vision-Language Models (VLMs) and video foundation models that enable machines to perceive, interpret, and generate rich representations of human behavior, expression, and interaction. Your research will span multi-modal reasoning, video understanding, and generative synthesis, enabling more natural and intuitive human-computer interaction at scale.

Responsibilities

Design and implement novel multi-modal architectures that fuse vision, language, and temporal signals for holistic human understanding
• Develop and train Vision-Language Models (VLMs) for tasks including visual question answering, image-text reasoning, and grounded human-centric understanding
• Build video foundation models capable of temporal reasoning, action synthesis and long-form video synthesis with applications to human behavior synthesis
• Research generative synthesis techniques for human-centric content including video generation, motion synthesis, and multi-modal content creation
• Conduct rigorous experiments to evaluate model performance across diverse benchmarks, analyze failure modes, and iterate on architectures to improve accuracy and generalization
• Contribute to the full research lifecycle from problem formulation and dataset curation through model development and evaluation

Minimum Qualifications
• Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta
• 2+ years of experience in multi-modal AI research, including hands-on work with Vision-Language Models, video understanding, or human-centric AI systems
• 2+ years of experience implementing and training large-scale neural networks using frameworks such as PyTorch, with experience on transformer-based architectures
• Experience designing and executing experiments to evaluate multi-modal model performance, including quantitative analysis across vision, language, and video benchmarks
• Experience writing production-quality or research-quality code in Python for multi-modal AI applications

Preferred Qualifications
• Experience developing or fine-tuning Vision-Language Models for human understanding tasks
• Experience with video foundation models, temporal transformers, or large-scale video pretraining
• Track record of contributing to published multi-modal AI research at venues such as CVPR, ICCV, or NeurIPS
• Experience with generative models for human synthesis including diffusion models, GANs, or autoregressive models for video or motion generation

Similar Jobs

More Jobs at Meta

More Consumer Technology Jobs

Find similar Research Scientist, Multi-Modal Human Understanding jobs: