V-JEPA 2
Meta’s open JEPA world model for video understanding, prediction and robot planning.
V-JEPA 2 is a self-supervised video world model from Meta FAIR built on the Joint-Embedding Predictive Architecture: instead of generating pixels it predicts representations of masked and future video. Pre-trained on over a million hours of internet video, it is used as a video encoder for understanding and retrieval; the action-conditioned V-JEPA 2-AC variant, fine-tuned on about 62 hours of robot video, plans zero-shot pick-and-place on real robot arms.
- Provider
- Meta
- Type
- World model
- Released
- Jun 11, 2025
- Price
- Open weights — self-host
- Input
- video, image
- Output
- embedding
- Open weights
- Yes
- License
- Apache 2.0
Best for
- Vision
- Agents
- Research
Strengths
- Open weights (Apache 2.0)
- Predicts in latent space, not pixels
- Zero-shot robot planning (V-JEPA 2-AC)
More from Meta
- Llama 4 Maverick — Meta. Meta’s 128-expert open multimodal MoE model.
- Llama 4 Scout — Meta. Efficient open model with an industry-leading 10M-token native context.
- Muse Spark 1.3 — Meta. Meta’s multimodal reasoning model for long-running agents and coding.
- Muse Image — Meta. Meta’s low-cost image model.
- Llama Guard 4 12B — Meta. Open multimodal safety classifier for prompts and responses.