VLA-JEPA: Teaching Robots How Actions Change the World
Zhongguancun Academy × USTC × SJTU × EIT propose VLA-JEPA, fusing VLA models and world models in latent space: assembly tasks completed with just 13 trajectories, 97.2% on LIBERO, shared by LeCun and Saining Xie.


VLA-JEPA: Teaching Robots How Actions Change the World
Zhongguancun Academy × USTC × SJTU × EIT propose VLA-JEPA, fusing VLA models and world models in latent space: assembly tasks completed with just 13 trajectories, 97.2% on LIBERO, shared by LeCun and Saining Xie.
Robot data is expensive to collect, limited in scale, and narrow in task coverage, while human videos and unlabeled manipulation videos on the internet are enormously abundant. How can vast amounts of human video help VLA models learn "how actions change the world"? VLA-JEPA, proposed by a team from USTC, Zhongguancun Academy, Shanghai Jiao Tong University, and the Eastern Institute of Technology, Ningbo, offers one answer: stop making the model chase "future frames" in pixel space, and instead follow Yann LeCun's JEPA line — learn and predict changes in world state in latent representation space.
As the first work combining VLA and world models ported to the lerobot framework, VLA-JEPA has been officially verified to complete simple assembly tasks with just 13 trajectories, and drew reposts and attention from LeCun and Saining Xie on social media.
Unifying human videos and robot demonstrations under a single "latent world model" training objective.
The One-Sentence Version
VLA-JEPA is a JEPA-style pretraining framework for VLA models. The current observation passes through the VLA backbone to yield latent action tokens; future frames provide supervision only through a target encoder, and the model must predict future states in latent space.
The design addresses the core bias of past latent-action pretraining: models tend to learn "pixel changes" rather than "action-induced state transitions." In internet video especially, camera movement, background changes, and irrelevant object motion can be more salient than the actual manipulation signal, causing latent actions to degenerate into compressed representations of the target image.
The human-video stage uses a latent world modeling alignment loss; the robot-data stage adds an action prediction loss on top.
Why This Approach Is Needed
Ideally, a latent action should capture "action-relevant state-transition semantics" — how the environment changes after an object is pushed, grasped, or moved — rather than simply recording which pixels changed. The VLA-JEPA paper identifies four widespread problems in existing latent-action pretraining:
- Pixel-level objectives bias representations toward appearance: texture, lighting, background, and viewpoint vary a lot and are easy to predict, yet correlate weakly with the degrees of freedom you actually want to control
- Real-world video amplifies irrelevant motion noise: camera shake and background movement overshadow the manipulation signal, and latent actions get dominated by noisy motion
- Information leakage degrades latent actions: when future observations participate in learning the action variable, the model can simply encode the future itself — semantically empty
- Multi-stage training pipelines are too complex: representation pretraining → latent action alignment → policy model, with objectives, data distributions, and representation spaces that are hard to keep consistent
The Method: Treat the Future as Supervision, Not Input
VLA-JEPA uses Qwen3-VL as the VLM backbone and introduces learnable latent action tokens to represent transitions between adjacent states. Video frames are mapped to world-state representations by a V-JEPA2 encoder; a predictor takes the current state and latent action to predict the future latent state, aligned with the future state produced by the target encoder.
On data with robot action annotations, a flow matching-based action head is attached to generate continuous end-effector trajectories.
The division of labor is clear:
- Human videos supply dynamic knowledge ("how actions change the world")
- Robot trajectories turn that dynamic knowledge into executable actions
Training is also more direct than a multi-stage latent-action pipeline: JEPA pretraining first, then fine-tune the action head.
Results
The paper evaluates on LIBERO, LIBERO-Plus, SimplerEnv, and real Franka tabletop manipulation tasks. Pretraining uses roughly 220,000 Something-Something-v2 human videos plus about 76,000 DROID robot demonstrations; LIBERO fine-tuning uses only about 2,000 simulated expert demonstrations; the real-world experiments use 100 demonstrations across three task types.
LIBERO / LIBERO-Plus
- LIBERO: average success rate 97.2%, with top scores on both the Object and LIBERO-10 suites. Notably, strong baselines like OpenVLA-OFT and pi0.5 rely on large amounts of robot data, while VLA-JEPA achieves comparable or better average performance with less training data
- LIBERO-Plus (multi-perturbation OOD): best results on 5 of 7 perturbation dimensions, average success rate 78.1%, clearly above OpenVLA-OFT (69.6%) and pi0-Fast (61.6)
This suggests the latent action learns not a single visual template but a representation closer to world-state changes — which is exactly where the robustness comes from.
SimplerEnv: Human Video Is Not a Panacea
SimplerEnv delivers a sobering reminder: on several visual matching tasks, the model scores even higher with human video removed. This shows VLA-JEPA's main value is not conjuring new action skills out of thin air, but strengthening robustness and stability on top of high-quality robot data.
Real Robot: The Post-Failure Re-Grasp
The real-world experiments use a Franka FR3 arm + Robotiq 2F-85 gripper + three D435 cameras. Compared with pi0 and pi0.5, VLA-JEPA shows an interesting phenomenon: after a failed first grasp, the model reopens the gripper and attempts a second grasp — a behavior the comparison models do not reliably exhibit.
The authors attribute this to "repeated-grasping knowledge" in human video — clips of humans adjusting and re-grasping after failure are common, while robot demonstration data rarely deliberately covers such recovery behavior. This is the most newsworthy part of the VLA-JEPA line: human video may not directly teach robot control, but it can supply the common sense of "how to recover."

How to Get It
The open-source resources are complete:
- arXiv: arxiv.org/abs/2602.10098
- Code: github.com/ginwind/VLA-JEPA
- Project page: ginwind.github.io/VLA-JEPA/
- Hugging Face: huggingface.co/ginwind/VLA-JEPA
Use Cases
- Robot policy research teams: match or beat strong baselines with less robot data, especially suited to long-tail tasks where data collection is costly
- VLA + world model research: the first such combination ported to the lerobot framework, a natural starting point for reproducing the VLA-JEPA line
- Industrial arms / tabletop manipulation: the post-failure re-grasp directly cuts failure rates in real production settings
- Human-video utilization research: the ablations clearly show human video "adds stability, not new skills" — a useful baseline for follow-up work
💡 Tip: VLA-JEPA does not prove human video can replace high-quality robot data; it clarifies the division of labor — robot data provides executable action grounding, human video provides broader world-dynamics experience. Treat human video as an "action-label substitute" and you will be disappointed; treat it as a "world-dynamics prior" and it genuinely pays off.
Project: github.com/ginwind/VLA-JEPA