Research Scientist, Multi-Modal Understanding & Synthesis

Meta

$150K — $180K *
Consumer Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or related field; PhD in a relevant area preferred.
  • 6+ years of AI research experience, particularly in multi-modal learning and generative models.
  • Proficient in Python and experienced with deep learning frameworks like PyTorch or TensorFlow.
  • Experience publishing research in reputable machine learning or AI forums.
  • Leadership experience in guiding research initiatives from concept to production.
  • Knowledge of multi-modal techniques such as vision-language integration.

Responsibilities

  • Lead original research in multi-modal learning with a focus on unifying understanding and generation.
  • Design and create world models that predict human behavior for simulation and reasoning.
  • Manage comprehensive research projects from formulation to real-time prototype integration.
  • Innovate synthesis approaches for cohesive generation across images, text, audio, and video.
  • Create evaluation frameworks and metrics to assess progress in multi-modal reasoning.
  • Mentor team members, offering guidance on multi-modal architectures and research methodologies.
  • Publish research findings in leading peer-reviewed conferences to contribute to the scientific community.

Benefits

  • Opportunity to work at the forefront of AI research and technology advancement.
  • Collaborative environment with access to world-class researchers and engineers.
  • Potential for publishing in top-tier AI venues, boosting professional visibility.
  • Engagement in impactful projects that drive next-gen AI products.
  • Career development through mentoring and leadership opportunities.
Full Job Description
In this role, you will advance the state of the art in building AI systems that perceive, reason across, and generate content spanning vision, language, audio, and other modalities. You will develop world models that learn rich internal representations of human behavior, enabling prediction, planning, and simulation. Collaborating with world-class researchers and engineers, you will define research directions, publish influential work, and translate breakthroughs into technologies that power Meta's next-generation AI products.

Responsibilities

Lead original research in multi-modal learning, developing architectures and algorithms that unify understanding and generation across vision, language, audio, and other modalities
• Design and build world models that learn predictive representations of human behavior, supporting capabilities such as simulation, planning, and reasoning
• Drive end-to-end research projects from problem formulation and dataset curation through model development, evaluation, and integration into real-time prototypes
• Develop novel approaches for multi-modal synthesis, enabling coherent generation of images, video, text, and audio from unified representations
• Establish rigorous evaluation frameworks, benchmarks, and metrics to measure progress in multi-modal reasoning and world modeling
• Mentor other engineers and researchers on the team, providing technical guidance on multi-modal architectures, generative models, and research best practices
• Publish research findings at top-tier peer-reviewed venues such as NeurIPS, ICLR, and CVPR to advance the broader scientific community

Minimum Qualifications
• Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
• PhD in Machine Learning, Computer Vision, Natural Language Processing, or a closely related field
• 6+ years of experience conducting AI research in multi-modal learning, generative models, or world models, including experience leading major research initiatives from conception through publication or production deployment
• Experience implementing and evaluating multi-modal systems using deep learning frameworks such as PyTorch or TensorFlow, with proficiency in Python
• Experience publishing original research in peer-reviewed machine learning or AI venues
• Experience driving cross-functional technical decisions and communicating research findings and trade-offs to both research and engineering audiences through written documents and presentations
• Experience with techniques spanning multiple modalities such as vision-language models, multi-modal transformers, or cross-modal representation learning

Preferred Qualifications
• Experience developing large-scale multi-modal foundation models or vision-language models
• Experience with world models, predictive learning, or model-based reinforcement learning for planning and reasoning
• First-author publications at top-tier venues such as NeurIPS, ICLR, or CVPR demonstrating contributions to multi-modal learning, generative models, or world models

Similar Jobs

More Jobs at Meta

More Consumer Technology Jobs

Find similar Research Scientist, Multi-Modal Understanding & Synthesis jobs: