V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models
TLDR
Introduces V-ReasonBench, a benchmark for evaluating video reasoning in generative models across four dimensions using synthetic and real-world sequences.
Reasoning
The paper presents a well-structured benchmark with clear dimensions and evaluation of six models, but its focus is on reasoning evaluation rather than world model development. Strengths include reproducibility and real-world data; weakness is limited novelty in world model concepts.
Read-first score
Read-first score 30.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 9.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 88.