V-JEPA 2

Meta’s open JEPA world model for video understanding, prediction and robot planning.

V-JEPA 2 is a self-supervised video world model from Meta FAIR built on the Joint-Embedding Predictive Architecture: instead of generating pixels it predicts representations of masked and future video. Pre-trained on over a million hours of internet video, it is used as a video encoder for understanding and retrieval; the action-conditioned V-JEPA 2-AC variant, fine-tuned on about 62 hours of robot video, plans zero-shot pick-and-place on real robot arms.

Provider
Meta
Type
World model
Released
Jun 11, 2025
Price
Open weights — self-host
Input
video, image
Output
embedding
Open weights
Yes
License
Apache 2.0

Best for

Strengths

More from Meta

Sources