Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

MBench: A Comprehensive Benchmark on Memory Capability for Video World Models

arXiv 2026 55.2 benchmark

TLDR

MBench is a benchmark evaluating memory capability of video world models via entity, environment, and causal consistency using real-captured videos.

Reasoning

The paper addresses a critical gap in evaluating long-term memory for video world models, with a systematic decomposition into three consistency dimensions and use of real-world data. However, it focuses narrowly on memory and does not cover other world model capabilities like interactivity or dynamics prediction.

Read-first score

Read-first score 55.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 41.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
91.9

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.91879498

Methodology quality 18%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation

Topical relevance 29%
58.6

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridge

Rank sensitivity

Stability: volatile; rank range: 408.

Keyword Scores

video world model
10
world model
9
generative world model
6
world dynamics prediction
6
world simulator
5
interactive world model
3
model-based reinforcement learning world model
2

Deep Analysis

Innovations

  • Proposes MBench, a comprehensive benchmark dedicated to evaluating memory capability of video world models.
  • Decomposes memory capability into three hierarchical dimensions: entity consistency, environment consistency, and causal consistency, further refined into 12 quantifiable sub-dimensions.
  • Uses rigorously curated real-captured long videos and rule-based quantitative matrices combined with VLM for objective and comprehensive consistency assessment.

Methodology

MBench systematically decomposes memory capability into three core dimensions (entity, environment, causal consistency) and 12 sub-dimensions. The benchmark is built upon curated real-captured long videos and evaluated using rule-based quantitative matrices and a Vision-Language Model (VLM) to enable objective consistency assessment.

Key Results

Extensive evaluations of mainstream state-of-the-art video world models reveal critical systemic limitations of existing methods in long-term state retention.

Tags

video world modelsmemory capabilitybenchmarkevaluationconsistencyCV