Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

From Generation to Simulation: How Far Are World Models from Being True Simulators?

arXiv 2026 54.2 survey, benchmark

TLDR

A capability-based survey assessing how far generative world models are from replacing traditional simulators, mapping 200 works across three technical routes.

Reasoning

The paper provides a systematic, structured analysis of world models against simulator capabilities, which is a valuable contribution. However, as a survey, it lacks original real-world experiments, and its conclusions rely on literature mapping rather than direct empirical validation.

Read-first score

Read-first score 54.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 57.

Methodology quality 18%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,evaluation,experiment,metric,validation

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Topical relevance 29%
81.4

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=github

Citation impact 18%
0

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. cited_by_count=0

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 770.

Keyword Scores

world model
10
world simulator
10
generative world model
10
video world model
8
interactive world model
7
world dynamics prediction
7
model-based reinforcement learning world model
5

Deep Analysis

Innovations

  • Capability-based yardstick with eight simulator capabilities (asset construction, physics engine, interaction, controllability, stability, state feedback, diversity, evaluation metrics) to assess generative world models
  • Systematic mapping of 200 representative works (2018–June 2026) across three technical routes (latent dynamics, video generation, joint-embedding prediction) onto the capability framework
  • Identification of state feedback as the most neglected cross-route shortcoming, with only 6 of 163 implementation papers exposing runtime state querying
  • Six research directions to bridge the gap between world models and traditional simulators: formalized physics, unified action interface, first-class state feedback, long-horizon stability, downstream-utility evaluation, and cross-route hybridization

Methodology

The study conducts a systematic literature review, collecting 200 representative works from 2018 to June 2026 along three technical routes. Each work is evaluated against an external yardstick of eight capabilities derived from traditional simulators, and the coverage and gaps are analyzed to characterize the remaining distance from generation to simulation.

Key Results

World models achieve functional substitution in interaction and controllability for specific scenarios but lack formal physical guarantees, structured state feedback, and reproducible long-horizon evolution; state feedback is especially neglected, with only 6 of 163 implementation papers providing a runtime interface for querying entity states or physical parameters.

Limitations

  • The analysis is based on a curated set of 200 representative works and may not capture all relevant approaches or very recent developments beyond June 2026.
  • The capability assessment relies on the authors' interpretation of published methods rather than direct empirical validation of each capability.
  • No new experimental validation of the identified gaps is provided; the study remains a survey and gap analysis.

Tags