How Should World Models Be Evaluated? A Decision-Making-Centric Position
TLDR
A position paper arguing that world models for decision-making should be evaluated on counterfactual reasoning and policy utility, not visual realism.
Reasoning
Strengths: Provides a clear L0-L7 evaluation ladder and addresses claim/evidence mismatch in current literature. Weaknesses: As a position paper, it lacks new empirical results or experiments to validate its proposed framework.
Read-first score
Read-first score 70.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 55.
Field roles
Rank sensitivity
Stability: volatile; rank range: 369.
Keyword Scores
Deep Analysis
Innovations
- L0–L7 ladder for categorizing world model evaluation from visual plausibility to policy optimization utility.
- Decision-making-centric evaluation framework emphasizing counterfactual reasoning, policy evaluation, planning, and policy optimization.
- Benchmark protocol focusing on counterfactual action fidelity, closed-loop rollout validity, reward/value prediction, policy-ranking agreement, optimization lift, model exploitability, and uncertainty calibration.
Methodology
The paper surveys the recent literature on world model evaluation, identifies claim/evidence mismatch, and proposes an L0–L7 ladder to organize evaluation levels. It then introduces a decision-making-centric evaluation framework and a benchmark protocol for embodied decision-making world models, without implementing or testing them empirically.
Key Results
The paper does not present experimental results; it is a position paper that argues for a shift in evaluation focus and provides a structured taxonomy and proposed evaluation criteria.
Limitations
- The paper is a position paper and does not provide empirical validation of the proposed framework or benchmark protocol.
- The L0–L7 ladder and evaluation criteria are based on the authors' interpretation and may not cover all aspects of world model evaluation.
- The proposed benchmark protocol is not implemented or tested, so its feasibility and effectiveness remain unverified.