AI JobFull-time
AI Researcher — Multimodal Models
Meta AI
Menlo Park, CA $250k–$400k Posted May 15, 2026
multimodal aivisionpytorchresearchllamapretraining
Job Description
Meta AI Research is seeking a researcher to advance multimodal foundation models. You will work on combining vision, audio, and language in unified model architectures, contributing to next-generation AI assistants and the Llama model family.
Responsibilities
- Design novel multimodal architectures combining vision and language
- Run large-scale pretraining and fine-tuning experiments
- Develop evaluation benchmarks for multimodal model capabilities
- Publish research at top venues (NeurIPS, ICML, CVPR, ICLR)
- Collaborate with product teams to apply research to Meta products
Requirements
- PhD in ML, Computer Vision, or NLP
- Track record of research publications
- Strong Python, PyTorch/JAX, and research engineering skills
- Experience with large-scale distributed training