Do World Action Models Generalize Better than VLAs? A Robustness Study
TLDR
Compares robustness of World Action Models (WAMs) vs. VLAs on robotic benchmarks under visual/language perturbations, finding WAMs more robust.
Reasoning
Strengths: Direct comparative study with explicit robustness evaluation on two benchmarks. Weaknesses: Abstract cut off, missing full results and limitations; no mention of real-world deployment beyond simulated benchmarks.
Read-first score
Read-first score 49.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 35.
Field roles
Rank sensitivity
Stability: volatile; rank range: 398.
Keyword Scores
Deep Analysis
Innovations
- Comparative robustness study of World Action Models (WAMs) versus Vision-Language-Action (VLA) models under visual and language perturbations.
- Demonstration that WAMs, leveraging dynamic prediction and spatiotemporal priors from web-scale video pretraining, achieve strong robustness (e.g., Cosmos-Policy 82.2% on LIBERO-Plus).
- Identification that hybrid approaches partially incorporating video-based dynamic learning exhibit intermediate robustness, highlighting the importance of how video priors are integrated.
Methodology
The study evaluates state-of-the-art VLA policies and recently released WAMs on the LIBERO-Plus and RoboTwin 2.0-Plus benchmarks under various visual and language perturbations. It compares performance metrics such as success rates across these conditions to assess robustness.
Key Results
WAMs achieve strong robustness, with LingBot-VA reaching 74.2% success rate on RoboTwin 2.0-Plus and Cosmos-Policy achieving 82.2% on LIBERO-Plus. VLAs like π0.5 can achieve comparable robustness on certain tasks but require extensive training with diverse robotic datasets and varied learning objectives.