Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Do World Action Models Generalize Better than VLAs? A Robustness Study

arXiv 26.3 2026 49.9 method, benchmark, application

TLDR

Compares robustness of World Action Models (WAMs) vs. VLAs on robotic benchmarks under visual/language perturbations, finding WAMs more robust.

Reasoning

Strengths: Direct comparative study with explicit robustness evaluation on two benchmarks. Weaknesses: Abstract cut off, missing full results and limitations; no mention of real-world deployment beyond simulated benchmarks.

Read-first score

Read-first score 49.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 35.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 18%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,dataset,metric,result

Topical relevance 29%
50

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,dataset

Citation impact 18%
40.2

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.40205754

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 398.

Keyword Scores

world model
9
world dynamics prediction
8
video world model
7
model-based reinforcement learning world model
4
generative world model
3
world simulator
2
interactive world model
2

Deep Analysis

Innovations

  • Comparative robustness study of World Action Models (WAMs) versus Vision-Language-Action (VLA) models under visual and language perturbations.
  • Demonstration that WAMs, leveraging dynamic prediction and spatiotemporal priors from web-scale video pretraining, achieve strong robustness (e.g., Cosmos-Policy 82.2% on LIBERO-Plus).
  • Identification that hybrid approaches partially incorporating video-based dynamic learning exhibit intermediate robustness, highlighting the importance of how video priors are integrated.

Methodology

The study evaluates state-of-the-art VLA policies and recently released WAMs on the LIBERO-Plus and RoboTwin 2.0-Plus benchmarks under various visual and language perturbations. It compares performance metrics such as success rates across these conditions to assess robustness.

Key Results

WAMs achieve strong robustness, with LingBot-VA reaching 74.2% success rate on RoboTwin 2.0-Plus and Cosmos-Policy achieving 82.2% on LIBERO-Plus. VLAs like π0.5 can achieve comparable robustness on certain tasks but require extensive training with diverse robotic datasets and varied learning objectives.

Tags