Predictive Latent World Modeling
for Vision-Language-Action Policies

arXiv 2026 / Preprint
Yu Liu1,2,* Hetian Guo1,* Tianlv Huang1 Ziyi Cai3 Wudi Chen1 Hantang Wang4 Qiutong Liu4
Yingzhi Peng5 Wei Han1 Peijun Tang2,† Jianan Wang2,† Zipei Fan1,‡ Zhiyuan Zha1 Xuan Song1
1JLU · 2Astribot · 3HITsz · 4PolyU · 5UTokyo
*Equal Contribution · Project Leads · Corresponding Author
PLaW-VLA overview teaser figure

PLaW-VLA learns predictive latent world representations from egocentric videos, robot trajectories, and simulation data. Future latent states inferred from history are incorporated into a unified VLA policy, helping the model anticipate scene evolution for long-horizon, fine-grained, and deformable-object manipulation.

Results

We show real-robot deployment videos and summarize real-world success rates for long-horizon, fine-grained, deformable-object, and zero-shot tasks.

I

Real-Robot Deployment

Complete deployment video for real-robot manipulation rollouts.

II

Quantitative Results

Success rates are reported over representative real-world manipulation tasks.

Real-world Robot Evaluation

Task Success on Real-World and Zero-Shot Manipulation

π0 π0.5 PLaW-VLA

* Zero-shot task variants

50 60 80
Fold Towel
47.5 42.5 50
Play Basketball
46.7 50 66.7
Pack Toy
43.3 41.7 55
Organize Stationery
30 50 80
Fold Towel*
35 40 48.3
Organize Stationery*

Method

PLaW-VLA factorizes policy generation into semantic grounding, latent-space prediction, and continuous action decoding.

PLaW-VLA framework: Mixture-of-Transformers backbone with VLM, World Model, and Action experts.
Framework. A unified Mixture-of-Transformers connects semantic grounding, predictive latent world modeling, and continuous action generation with structured causal attention.
VLM expert Grounds language, observation, and proprioception into task context.
World model expert Future-state queries attend to observation history to predict future latent states.
Action expert Decodes continuous action chunks conditioned on predicted future latents.

BibTeX

If you find PLaW-VLA useful, please cite our paper.

@article{liu2026plawvla,
  title   = {PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies},
  author  = {Liu, Yu and Guo, Hetian and Huang, Tianlv and Cai, Ziyi and Chen, Wudi and
             Wang, Hantang and Liu, Qiutong and Peng, Yingzhi and Han, Wei and
             Tang, Peijun and Wang, Jianan and Fan, Zipei and Zha, Zhiyuan and Song, Xuan},
  journal = {arXiv preprint arXiv:XXXX.XXXXX},
  year    = {2026}
}

Acknowledgements

We gratefully acknowledge the Physical Intelligence team, whose open-source contributions have provided a valuable foundation for this work. We also thank Hongji Huang for filming and producing the real-robot experiment videos.