Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation

arXiv 2026 68.2 benchmark

TLDR

WBench is a multi-turn benchmark evaluating interactive video world models across five dimensions with 289 test cases and 22 automatic metrics.

Reasoning

The paper addresses a clear gap with a comprehensive benchmark, validated against human judgments, and evaluates 20 models. However, the abstract does not detail the specific models or results, and the benchmark's generalizability may be limited by its predefined interaction types.

Read-first score

Read-first score 68.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 49.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
85.7

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.85730805

Reproducibility 18%
81

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=code,github

Topical relevance 29%
70

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 18%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,metric

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchorReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 418.

Keyword Scores

world model
10
interactive world model
10
video world model
9
world dynamics prediction
7
world simulator
6
generative world model
5
model-based reinforcement learning world model
2

Deep Analysis

Innovations

  • Comprehensive multi-turn benchmark for interactive world model evaluation
  • Five evaluation dimensions: video quality, setting adherence, interaction adherence, consistency, and physics compliance
  • 289 test cases and 1,058 interaction turns covering diverse scenes, styles, subjects, and perspectives
  • Four interaction types: navigation, subject action, event editing, and perspective switching
  • Unified navigation control supporting text, 6-DoF pose, and discrete-action inputs
  • 22 automatic sub-metrics combining specialist vision models and large multimodal models, validated against human judgments
  • Diagnostic insights across 20 state-of-the-art models

Methodology

WBench is constructed with 289 test cases, each specifying a world setting and a multi-turn interaction sequence, covering diverse scenes, styles, subjects, first- and third-person perspectives, and four interaction types (navigation, subject action, event editing, perspective switching). Evaluation uses 22 automatic sub-metrics that combine specialist vision models with large multimodal models, all validated against human judgments. The benchmark is applied to 20 state-of-the-art models.

Key Results

Across 20 state-of-the-art models, no single model performs strongly across all five dimensions. The study provides detailed diagnostic insights into the characteristic strengths, weaknesses, and open challenges of each model.

Tags

interactive world modelsbenchmarkmulti-turn evaluationvideo qualityphysics complianceinteraction adherenceCV