Research Scientist, Multi-Modal Understanding & Synthesis
Meta - Redmond, WA
Hiring: Research Scientist, Multi-Modal Understanding & Synthesis Company: Meta Location: Redmond, WA Job Posted Time: 2026-09-16 15:58:56 Employment Type: Full-time Target Skills & Keywords: Multi-Modal Learning, World Models, Generative Models, Vision-Language Models, Multi-Modal Transformers, Cross-Modal Representation Learning, Deep Learning, PyTorch, TensorFlow, Python, Model-Based Reinforcement Learning, Research Publications Experience: - 6+ years of experience conducting AI research in multi-modal learning, generative models, or world models - Experience leading major research initiatives from conception through publication or production deployment - Experience driving cross-functional technical decisions and communicating research findings and trade-offs to research and engineering audiences Required Skills: - Lead original research in multi-modal learning, developing architectures and algorithms that unify understanding and generation across vision, language, audio, and other modalities - Design and build world models that learn predictive representations of human behavior, supporting simulation, planning, and reasoning - Drive end-to-end research projects from problem formulation and dataset curation through model development, evaluation, and integration into real-time prototypes - Develop novel approaches for multi-modal synthesis, enabling coherent generation of images, video, text, and audio from unified representations - Establish rigorous evaluation frameworks, benchmarks, and metrics to measure progress in multi-modal reasoning and world modeling - Mentor engineers and researchers on multi-modal architectures, generative models, and research best practices - Publish research findings at top-tier peer-reviewed venues such as NeurIPS, ICLR, and CVPR - Implement and evaluate multi-modal systems using deep learning frameworks such as PyTorch or TensorFlow, with proficiency in Python - Apply techniques spanning multiple modalities such as vision-language models, multi-modal transformers, or cross-modal representation learning Qualifications: - Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience - PhD in Machine Learning, Computer Vision, Natural Language Processing, or a closely related field - Experience publishing original research in peer-reviewed machine learning or AI venues - Preferred: Experience developing large-scale multi-modal foundation models or vision-language models - Preferred: Experience with world models, predictive learning, or model-based reinforcement learning for planning and reasoning - Preferred: First-author publications at top-tier venues such as NeurIPS, ICLR, or CVPR Compensation: $184,000.00/yr - $257,000.00/yr Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don’t miss this opportunity to join a forward-thinking team!