WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors
TLDR
Introduces WorldReasonBench, a benchmark to test video generators as world-state predictors, revealing a gap between visual plausibility and reasoning.
Reasoning
Strengths: Novel benchmark with structured QA and human-aligned evaluation, covering multiple reasoning dimensions. Weaknesses: Limited to 436 test cases; no discussion of scalability or generalization to diverse domains.
Read-first score
Read-first score 65.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 47.
Field roles
Rank sensitivity
Stability: volatile; rank range: 429.
Keyword Scores
Deep Analysis
Innovations
- WorldReasonBench: a benchmark that reframes video generation evaluation as world-state prediction, testing physical, social, logical, and informational consistency.
- Two-part human-aligned evaluation methodology: Process-aware Reasoning Verification (structured QA and reasoning-phase diagnostics) and Multi-dimensional Quality Assessment (reasoning quality, temporal consistency, visual aesthetics).
- WorldRewardBench: a preference benchmark with approximately 6K expert-annotated pairs over 1.4K videos for pair-wise and point-wise reward-model evaluation.
- Identification of a persistent gap between visual plausibility and world reasoning in modern video generators.
Methodology
WorldReasonBench contains 436 curated test cases with structured ground-truth QA annotations spanning four reasoning dimensions (physical, social, logical, informational) and 22 subcategories. Evaluation uses a two-part methodology: Process-aware Reasoning Verification detects temporal and causal failures via structured QA and reasoning-phase diagnostics, while Multi-dimensional Quality Assessment scores reasoning quality, temporal consistency, and visual aesthetics for ranking and reward modeling. Additionally, WorldRewardBench provides approximately 6K expert-annotated preference pairs over 1.4K videos to support reward-model evaluation.
Key Results
Across modern video generators, results expose a persistent gap between visual plausibility and world reasoning: videos can look convincing while failing dynamics, causality, or information preservation.
Limitations
- Benchmark size is limited to 436 test cases, which may not capture the full diversity of real-world scenarios.
- Evaluation relies on structured ground-truth QA annotations, which may introduce subjectivity and are costly to scale.
- The benchmark covers only four reasoning dimensions (physical, social, logical, informational) and 22 subcategories, potentially missing other aspects of world simulation.