Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

How Should World Models Be Evaluated? A Decision-Making-Centric Position

arXiv 2026 70.1 survey

TLDR

A position paper arguing that world models for decision-making should be evaluated on counterfactual reasoning and policy utility, not visual realism.

Reasoning

Strengths: Provides a clear L0-L7 evaluation ladder and addresses claim/evidence mismatch in current literature. Weaknesses: As a position paper, it lacks new empirical results or experiments to validate its proposed framework.

Read-first score

Read-first score 70.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 55.

Methodology quality 18%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,experiment,metric,result,validation

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
94.8

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.94785012

Topical relevance 29%
78.6

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=artifact

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 369.

Keyword Scores

world model
10
model-based reinforcement learning world model
9
interactive world model
8
world dynamics prediction
8
world simulator
7
video world model
7
generative world model
6

Deep Analysis

Innovations

  • L0–L7 ladder for categorizing world model evaluation from visual plausibility to policy optimization utility.
  • Decision-making-centric evaluation framework emphasizing counterfactual reasoning, policy evaluation, planning, and policy optimization.
  • Benchmark protocol focusing on counterfactual action fidelity, closed-loop rollout validity, reward/value prediction, policy-ranking agreement, optimization lift, model exploitability, and uncertainty calibration.

Methodology

The paper surveys the recent literature on world model evaluation, identifies claim/evidence mismatch, and proposes an L0–L7 ladder to organize evaluation levels. It then introduces a decision-making-centric evaluation framework and a benchmark protocol for embodied decision-making world models, without implementing or testing them empirically.

Key Results

The paper does not present experimental results; it is a position paper that argues for a shift in evaluation focus and provides a structured taxonomy and proposed evaluation criteria.

Limitations

  • The paper is a position paper and does not provide empirical validation of the proposed framework or benchmark protocol.
  • The L0–L7 ladder and evaluation criteria are based on the authors' interpretation and may not cover all aspects of world model evaluation.
  • The proposed benchmark protocol is not implemented or tested, so its feasibility and effectiveness remain unverified.

Tags

world modelsevaluationdecision-makingAImetricsposition paperLG