τ₀-VLA: Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
The τ₀-VLA authors introduce world-model-guided test-time computation at the high level, allowing the model to search over subtask alternatives before committing. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training, and the authors report that allocating additional test-time computation improves next-subtask prediction accuracy in both in-domain and distribution-shifted settings, with those gains carrying into higher closed-loop task success [ref:ref-5].