Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

World Action Models: A Survey

arXiv 2026 65.9 survey

TLDR

A survey clarifying World Action Models as embodied predictive-action models distinct from video generators, with a taxonomy and analysis of design trade-offs.

Reasoning

Strengths include a clear taxonomy and unified discussion of key properties like interactability and causality. Weaknesses: as a survey, it lacks novel experiments or empirical contributions, and the abstract does not detail specific real-world evaluations.

Read-first score

Read-first score 65.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 58.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
94.2

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.94211446

Topical relevance 29%
82.9

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 18%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,evaluation

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=github

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 438.

Keyword Scores

world model
10
generative world model
9
video world model
9
world dynamics prediction
9
interactive world model
8
world simulator
7
model-based reinforcement learning world model
6

Deep Analysis

Innovations

  • Clarifies the boundaries between World Action Models, video generation models, action-grounded video world models, Vision-Language-Action policies, and broad world models.
  • Organizes existing works through two complementary views: what each method generates (rendered futures, latent futures, video-generation-free action reasoning) and decomposition by predictive substrate, backbone, action coupling, and deployment regime.
  • Provides a unified discussion of interactability, causality, persistence, physical plausibility, and generalization across methods.
  • Identifies a consistent design pattern: WAMs are predictive-action methods that trade representational richness against compute, memory, latency, and action-label cost.
  • Highlights a field trend toward methods that generate less of the future while preserving what control requires.

Methodology

The survey conducts a comprehensive literature review of World Action Models, categorizing them based on two complementary views: the type of future they generate (rendered, latent, or action reasoning without video generation) and their decomposition by predictive substrate, backbone, action coupling, and deployment regime. It then discusses key properties such as interactability, causality, persistence, physical plausibility, and generalization, followed by an analysis of data, evaluation, and open challenges.

Key Results

The survey identifies a consistent design pattern where World Action Models are not simply video generators with action heads but predictive-action methods that trade representational richness against compute, memory, latency, and action-label cost. It observes a field trend toward methods that generate less of the future while preserving what control requires.

Limitations

  • The rapid expansion of the field may make the survey's categorization incomplete or quickly outdated.
  • The proposed taxonomy and boundaries, while clarifying, may still involve subjective judgments due to the blurred nature of the field.

Tags

world action modelsembodied AIvideo generationworld modelssurveyROCV