Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

A Mechanistic View on Video Generation as World Models: State and Dynamics

arXiv 26.1 2026 70.2 survey, theory

TLDR

Proposes a taxonomy bridging video generation and world models via state construction and dynamics modeling, advocating functional benchmarks.

Reasoning

Strengths: Clear taxonomy linking video generation to world model theory, identifies key frontiers. Weaknesses: Conceptual paper without empirical validation or real-world experiments.

Read-first score

Read-first score 70.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 46.

Methodology quality 25%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,benchmark,dataset,evaluation,experiment,metric

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Topical relevance 42%
65.7

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 112.

Keyword Scores

world model
9
generative world model
9
video world model
9
world simulator
8
world dynamics prediction
8
interactive world model
2
model-based reinforcement learning world model
1

Deep Analysis

Innovations

  • Proposes a novel taxonomy for video generation as world models centered on State Construction and Dynamics Modeling.
  • Categorizes state construction into implicit (context management) and explicit (latent compression) paradigms.
  • Analyzes dynamics modeling through knowledge integration and architectural reformulation.
  • Advocates for a transition in evaluation from visual fidelity to functional benchmarks testing physical persistence and causal reasoning.
  • Identifies two critical frontiers: enhancing persistence via data-driven memory and compressed fidelity, and advancing causality through latent factor decoupling and reasoning-prior integration.

Methodology

The paper presents a conceptual analysis and taxonomy based on reviewing existing video generation models and world model theories. It categorizes approaches into implicit and explicit state construction, and dynamics modeling via knowledge integration and architectural reformulation. No new experiments or datasets are introduced.

Key Results

The paper proposes a taxonomy of state construction (implicit/explicit) and dynamics modeling (knowledge integration, architectural reformulation), and advocates for functional benchmarks. It identifies two critical frontiers: enhancing persistence and advancing causality.

Limitations

  • Current video generation models are stateless and lack explicit state construction, creating a gap with classic world model theories.
  • Evaluation metrics focus on visual fidelity rather than physical persistence and causal reasoning.
  • Existing models face challenges in long-term persistence and causal reasoning, requiring data-driven memory and latent factor decoupling.

Tags