From Generation to Simulation: How Far Are World Models from Being True Simulators?
TLDR
A capability-based survey assessing how far generative world models are from replacing traditional simulators, mapping 200 works across three technical routes.
Reasoning
The paper provides a systematic, structured analysis of world models against simulator capabilities, which is a valuable contribution. However, as a survey, it lacks original real-world experiments, and its conclusions rely on literature mapping rather than direct empirical validation.
Read-first score
Read-first score 54.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 57.
Field roles
Rank sensitivity
Stability: volatile; rank range: 770.
Keyword Scores
Deep Analysis
Innovations
- Capability-based yardstick with eight simulator capabilities (asset construction, physics engine, interaction, controllability, stability, state feedback, diversity, evaluation metrics) to assess generative world models
- Systematic mapping of 200 representative works (2018–June 2026) across three technical routes (latent dynamics, video generation, joint-embedding prediction) onto the capability framework
- Identification of state feedback as the most neglected cross-route shortcoming, with only 6 of 163 implementation papers exposing runtime state querying
- Six research directions to bridge the gap between world models and traditional simulators: formalized physics, unified action interface, first-class state feedback, long-horizon stability, downstream-utility evaluation, and cross-route hybridization
Methodology
The study conducts a systematic literature review, collecting 200 representative works from 2018 to June 2026 along three technical routes. Each work is evaluated against an external yardstick of eight capabilities derived from traditional simulators, and the coverage and gaps are analyzed to characterize the remaining distance from generation to simulation.
Key Results
World models achieve functional substitution in interaction and controllability for specific scenarios but lack formal physical guarantees, structured state feedback, and reproducible long-horizon evolution; state feedback is especially neglected, with only 6 of 163 implementation papers providing a runtime interface for querying entity states or physical parameters.
Limitations
- The analysis is based on a curated set of 200 representative works and may not capture all relevant approaches or very recent developments beyond June 2026.
- The capability assessment relies on the authors' interpretation of published methods rather than direct empirical validation of each capability.
- No new experimental validation of the identified gaps is provided; the study remains a survey and gap analysis.