Current World Models Lack a Persistent State Core
TLDR
Paper introduces WRBench to test if world models maintain persistent state when unobserved, finding current models fail to evolve events during occlusion.
Reasoning
Strengths: novel benchmark, systematic evaluation across 23 models and 9600 videos, identifies a fundamental limitation. Weaknesses: no real-world validation, limited to synthetic video models.
Read-first score
Read-first score 64.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 54.
Field roles
Rank sensitivity
Stability: volatile; rank range: 429.
Keyword Scores
Deep Analysis
Innovations
- WRBench: first systematic diagnostic benchmark for persistent state core in world models
- Evaluation chain treating camera motion as an intervention on observability
- Identification that current world models lack persistent state core and resume abandoned states
Methodology
WRBench is a diagnostic benchmark that treats camera motion as an intervention on observability. It evaluates world models through a human-calibrated chain: whether the camera executes the requested interaction, whether the scene stays continuous and identifiable while in view, and whether a returning target remains consistent with the event that was set in motion. The benchmark uses 9,600 videos from 23 models across four control paradigms.
Key Results
Across all 23 models and 4 control paradigms, current world models fail to advance the world state when unobserved; they resume a returning target in the state at which it was abandoned rather than evolving the event while unseen.
Limitations
- The benchmark does not propose a method to achieve persistent state core.
- The study is limited to evaluating existing models; no new model or training paradigm is introduced.