Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Current World Models Lack a Persistent State Core

arXiv 2026 64.7 benchmark

TLDR

Paper introduces WRBench to test if world models maintain persistent state when unobserved, finding current models fail to evolve events during occlusion.

Reasoning

Strengths: novel benchmark, systematic evaluation across 23 models and 9600 videos, identifies a fundamental limitation. Weaknesses: no real-world validation, limited to synthetic video models.

Read-first score

Read-first score 64.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 54.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
94.7

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.94670088

Methodology quality 18%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,metric

Topical relevance 29%
77.1

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 429.

Keyword Scores

world model
10
generative world model
9
world dynamics prediction
9
world simulator
8
video world model
8
interactive world model
7
model-based reinforcement learning world model
3

Deep Analysis

Innovations

  • WRBench: first systematic diagnostic benchmark for persistent state core in world models
  • Evaluation chain treating camera motion as an intervention on observability
  • Identification that current world models lack persistent state core and resume abandoned states

Methodology

WRBench is a diagnostic benchmark that treats camera motion as an intervention on observability. It evaluates world models through a human-calibrated chain: whether the camera executes the requested interaction, whether the scene stays continuous and identifiable while in view, and whether a returning target remains consistent with the event that was set in motion. The benchmark uses 9,600 videos from 23 models across four control paradigms.

Key Results

Across all 23 models and 4 control paradigms, current world models fail to advance the world state when unobserved; they resume a returning target in the state at which it was abandoned rather than evolving the event while unseen.

Limitations

  • The benchmark does not propose a method to achieve persistent state core.
  • The study is limited to evaluating existing models; no new model or training paradigm is introduced.

Tags

world modelsbenchmarkpersistent statecamera motionobservabilityCV