Meta is seeking a Research Scientist to advance multi-modal AI technologies for human understanding and synthesis. In this role, you will develop Vision-Language Models (VLMs) and video foundation models that enable machines to perceive, interpret, and generate rich representations of human behavior, expression, and interaction. Your research will span multi-modal reasoning, video understanding, and generative synthesis, enabling more natural and intuitive human-computer interaction at scale.
Responsibilities
Design and implement novel multi-modal architectures that fuse vision, language, and temporal signals for holistic human understanding
• Develop and train Vision-Language Models (VLMs) for tasks including visual question answering, image-text reasoning, and grounded human-centric understanding
• Build video foundation models capable of temporal reasoning, action synthesis and long-form video synthesis with applications to human behavior synthesis
• Research generative synthesis techniques for human-centric content including video generation, motion synthesis, and multi-modal content creation
• Conduct rigorous experiments to evaluate model performance across diverse benchmarks, analyze failure modes, and iterate on architectures to improve accuracy and generalization
• Contribute to the full research lifecycle from problem formulation and dataset curation through model development and evaluation
Minimum Qualifications
• Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta
• 2+ years of experience in multi-modal AI research, including hands-on work with Vision-Language Models, video understanding, or human-centric AI systems
• 2+ years of experience implementing and training large-scale neural networks using frameworks such as PyTorch, with experience on transformer-based architectures
• Experience designing and executing experiments to evaluate multi-modal model performance, including quantitative analysis across vision, language, and video benchmarks
• Experience writing production-quality or research-quality code in Python for multi-modal AI applications
Preferred Qualifications
• Experience developing or fine-tuning Vision-Language Models for human understanding tasks
• Experience with video foundation models, temporal transformers, or large-scale video pretraining
• Track record of contributing to published multi-modal AI research at venues such as CVPR, ICCV, or NeurIPS
• Experience with generative models for human synthesis including diffusion models, GANs, or autoregressive models for video or motion generation