Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

arXiv 2026 65.9 benchmark

TLDR

Introduces WorldReasonBench, a benchmark to test video generators as world-state predictors, revealing a gap between visual plausibility and reasoning.

Reasoning

Strengths: Novel benchmark with structured QA and human-aligned evaluation, covering multiple reasoning dimensions. Weaknesses: Limited to 436 test cases; no discussion of scalability or generalization to diverse domains.

Read-first score

Read-first score 65.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 47.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 18%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,result

Citation impact 18%
75.1

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.75056323

Reproducibility 18%
73

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=github

Topical relevance 29%
67.1

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchorReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 429.

Keyword Scores

video world model
9
world dynamics prediction
9
world simulator
8
world model
7
generative world model
7
interactive world model
6
model-based reinforcement learning world model
1

Deep Analysis

Innovations

  • WorldReasonBench: a benchmark that reframes video generation evaluation as world-state prediction, testing physical, social, logical, and informational consistency.
  • Two-part human-aligned evaluation methodology: Process-aware Reasoning Verification (structured QA and reasoning-phase diagnostics) and Multi-dimensional Quality Assessment (reasoning quality, temporal consistency, visual aesthetics).
  • WorldRewardBench: a preference benchmark with approximately 6K expert-annotated pairs over 1.4K videos for pair-wise and point-wise reward-model evaluation.
  • Identification of a persistent gap between visual plausibility and world reasoning in modern video generators.

Methodology

WorldReasonBench contains 436 curated test cases with structured ground-truth QA annotations spanning four reasoning dimensions (physical, social, logical, informational) and 22 subcategories. Evaluation uses a two-part methodology: Process-aware Reasoning Verification detects temporal and causal failures via structured QA and reasoning-phase diagnostics, while Multi-dimensional Quality Assessment scores reasoning quality, temporal consistency, and visual aesthetics for ranking and reward modeling. Additionally, WorldRewardBench provides approximately 6K expert-annotated preference pairs over 1.4K videos to support reward-model evaluation.

Key Results

Across modern video generators, results expose a persistent gap between visual plausibility and world reasoning: videos can look convincing while failing dynamics, causality, or information preservation.

Limitations

  • Benchmark size is limited to 436 test cases, which may not capture the full diversity of real-world scenarios.
  • Evaluation relies on structured ground-truth QA annotations, which may introduce subjectivity and are costly to scale.
  • The benchmark covers only four reasoning dimensions (physical, social, logical, informational) and 22 subcategories, potentially missing other aspects of world simulation.

Tags

video generationbenchmarkworld-state predictionreasoningconsistencyevaluationCV