We show real-robot deployment videos and summarize real-world success rates for long-horizon, fine-grained, deformable-object, and zero-shot tasks.
Complete deployment video for real-robot manipulation rollouts.
Success rates are reported over representative real-world manipulation tasks.
PLaW-VLA factorizes policy generation into semantic grounding, latent-space prediction, and continuous action decoding.
If you find PLaW-VLA useful, please cite our paper.
@article{liu2026plawvla,
title = {PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies},
author = {Liu, Yu and Guo, Hetian and Huang, Tianlv and Cai, Ziyi and Chen, Wudi and
Wang, Hantang and Liu, Qiutong and Peng, Yingzhi and Han, Wei and
Tang, Peijun and Wang, Jianan and Fan, Zipei and Zha, Zhiyuan and Song, Xuan},
journal = {arXiv preprint arXiv:XXXX.XXXXX},
year = {2026}
}
We gratefully acknowledge the Physical Intelligence team, whose open-source contributions have provided a valuable foundation for this work. We also thank Hongji Huang for filming and producing the real-robot experiment videos.