Paper Detail
Xiaowei Cai, Yunuo Cai, Bingao Chen, Jingxiao Chen, Zhi Chen, Siyuan Feng, Tengyu Hou, Jingshun Huang, Han Jiang, Runkun Ju, Dong Li, Mingxiang Li, Shaowei Li, Xinchen Li, Yifan Li, Yi Liu, Zhongyuan Liu, Jianlan Luo, Junwen Miao, Ruiqi Ni, Buqing Nie, Mingjie Pan, Xinlin Ren, Jianheng Song, Jiaxu Wang, Peiqi Wang, Sen Wang, Xiaoyan Wang, Dafeng Wei, Dongming Wu, Pengwei Xie, Pu Yang, Hangjian Ye, Xiangyu Yue, Jinyu Zhang, Qinglin Zhang, Xueyong Zhao, Pengfei Zhou, Yue Zhou
Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce $τ_0$-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.
No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.
No ranking explanation is available yet.
No tags.
@article{cai20260,
title = {\$τ\_0\$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation},
author = {Xiaowei Cai and Yunuo Cai and Bingao Chen and Jingxiao Chen and Zhi Chen and Siyuan Feng and Tengyu Hou and Jingshun Huang and Han Jiang and Runkun Ju and Dong Li and Mingxiang Li and Shaowei Li and Xinchen Li and Yifan Li and Yi Liu and Zhongyuan Liu and Jianlan Luo and Junwen Miao and Ruiqi Ni and Buqing Nie and Mingjie Pan and Xinlin Ren and Jianheng Song and Jiaxu Wang and Peiqi Wang and Sen Wang and Xiaoyan Wang and Dafeng Wei and Dongming Wu and Pengwei Xie and Pu Yang and Hangjian Ye and Xiangyu Yue and Jinyu Zhang and Qinglin Zhang and Xueyong Zhao and Pengfei Zhou and Yue Zhou},
year = {2026},
abstract = {Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce \$τ\_0\$-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through w},
url = {https://arxiv.org/abs/2608.16885},
keywords = {cs.RO},
eprint = {2608.16885},
archiveprefix = {arXiv},
}
{}