VideoVerse: How Far is Your T2V Generator from a World Model?
TLDR
VideoVerse benchmark evaluates T2V models on temporal causality and world knowledge, revealing gaps in world model capabilities.
Reasoning
Strengths include a comprehensive benchmark with 300 prompts and 793 evaluation questions, using human-aligned QA to assess world knowledge and temporal causality. Weaknesses are the focus solely on text-to-video generation, lacking interactive or RL-based world model evaluation.
Read-first score
Read-first score 38.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 37.
Field roles
Rank sensitivity
Stability: volatile; rank range: 207.
Keyword Scores
Deep Analysis
Innovations
- Introduction of VideoVerse benchmark that evaluates T2V models on complex temporal causality and world knowledge
- Design of ten evaluation dimensions covering dynamic and static properties
- Development of a human preference-aligned QA-based evaluation pipeline using modern vision-language models
Methodology
VideoVerse collects representative videos across diverse domains, extracts event-level descriptions with inherent temporal causality, and rewrites them into text-to-video prompts by independent annotators. For each prompt, ten evaluation dimensions covering dynamic and static properties are designed, resulting in 300 prompts, 815 events, and 793 evaluation questions. A human preference-aligned QA-based evaluation pipeline using modern vision-language models is developed to systematically benchmark leading open- and closed-source T2V systems.
Key Results
The benchmark reveals a significant gap between current T2V models and desired world modeling abilities, particularly in understanding complex temporal causality and world knowledge.