Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Simulating the Visual World with Artificial Intelligence: A Roadmap

arXiv 25.11 2025 65.1 survey

TLDR

Survey conceptualizing video foundation models as implicit world models with a video renderer, tracing four generations toward interactive, physically plausible simulation.

Reasoning

Strengths: Provides a clear conceptual framework linking video generation to world models, with a structured roadmap of four generations. Weaknesses: As a survey, it lacks new empirical experiments or real-world validation, and the abstract does not detail specific benchmarks or datasets.

Read-first score

Read-first score 65.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 57.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Topical relevance 42%
81.4

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 25%
50

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=none

Reproducibility 25%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,github

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 335.

Keyword Scores

world model
10
world simulator
9
generative world model
9
video world model
9
interactive world model
8
world dynamics prediction
8
model-based reinforcement learning world model
4

Deep Analysis

Innovations

  • Conceptual framework of video foundation models as combination of an implicit world model and a video renderer
  • Four-generation progression of video generation towards world models with increasing capabilities
  • Roadmap and design principles for next-generation world models, including the role of agent intelligence

Methodology

This survey provides a systematic overview of the evolution of video generation, categorizing modern video foundation models into two core components: an implicit world model and a video renderer. It traces the progression through four generations, defining core characteristics, highlighting representative works, and examining application domains such as robotics, autonomous driving, and interactive gaming.

Key Results

The survey traces the progression of video generation through four generations, culminating in a world model that embodies intrinsic physical plausibility, real-time multimodal interaction, and planning capabilities across multiple spatiotemporal scales.

Limitations

  • Current video generation models lack intrinsic physical plausibility and real-time multimodal interaction capabilities
  • Insufficient planning capabilities across multiple spatiotemporal scales in existing models
  • Open challenges remain in designing next-generation world models, including the role of agent intelligence in shaping and evaluating these systems

Tags