WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
TLDR
WBench is a multi-turn benchmark evaluating interactive video world models across five dimensions with 289 test cases and 22 automatic metrics.
Reasoning
The paper addresses a clear gap with a comprehensive benchmark, validated against human judgments, and evaluates 20 models. However, the abstract does not detail the specific models or results, and the benchmark's generalizability may be limited by its predefined interaction types.
Read-first score
Read-first score 68.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 49.
Field roles
Rank sensitivity
Stability: volatile; rank range: 418.
Keyword Scores
Deep Analysis
Innovations
- Comprehensive multi-turn benchmark for interactive world model evaluation
- Five evaluation dimensions: video quality, setting adherence, interaction adherence, consistency, and physics compliance
- 289 test cases and 1,058 interaction turns covering diverse scenes, styles, subjects, and perspectives
- Four interaction types: navigation, subject action, event editing, and perspective switching
- Unified navigation control supporting text, 6-DoF pose, and discrete-action inputs
- 22 automatic sub-metrics combining specialist vision models and large multimodal models, validated against human judgments
- Diagnostic insights across 20 state-of-the-art models
Methodology
WBench is constructed with 289 test cases, each specifying a world setting and a multi-turn interaction sequence, covering diverse scenes, styles, subjects, first- and third-person perspectives, and four interaction types (navigation, subject action, event editing, perspective switching). Evaluation uses 22 automatic sub-metrics that combine specialist vision models with large multimodal models, all validated against human judgments. The benchmark is applied to 20 state-of-the-art models.
Key Results
Across 20 state-of-the-art models, no single model performs strongly across all five dimensions. The study provides detailed diagnostic insights into the characteristic strengths, weaknesses, and open challenges of each model.