MBench: A Comprehensive Benchmark on Memory Capability for Video World Models
TLDR
MBench is a benchmark evaluating memory capability of video world models via entity, environment, and causal consistency using real-captured videos.
Reasoning
The paper addresses a critical gap in evaluating long-term memory for video world models, with a systematic decomposition into three consistency dimensions and use of real-world data. However, it focuses narrowly on memory and does not cover other world model capabilities like interactivity or dynamics prediction.
Read-first score
Read-first score 55.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 41.
Field roles
Rank sensitivity
Stability: volatile; rank range: 408.
Keyword Scores
Deep Analysis
Innovations
- Proposes MBench, a comprehensive benchmark dedicated to evaluating memory capability of video world models.
- Decomposes memory capability into three hierarchical dimensions: entity consistency, environment consistency, and causal consistency, further refined into 12 quantifiable sub-dimensions.
- Uses rigorously curated real-captured long videos and rule-based quantitative matrices combined with VLM for objective and comprehensive consistency assessment.
Methodology
MBench systematically decomposes memory capability into three core dimensions (entity, environment, causal consistency) and 12 sub-dimensions. The benchmark is built upon curated real-captured long videos and evaluated using rule-based quantitative matrices and a Vision-Language Model (VLM) to enable objective consistency assessment.
Key Results
Extensive evaluations of mainstream state-of-the-art video world models reveal critical systemic limitations of existing methods in long-term state retention.