Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

DeepVerse: 4D Autoregressive Video Generation as a World Model

arXiv 25.6 2025 60.3 method

TLDR

DeepVerse is a 4D autoregressive video generation world model that incorporates explicit geometric predictions to reduce drift and enhance temporal consistency.

Reasoning

The paper introduces a novel approach to world modeling by integrating geometric constraints, which addresses a key limitation of existing models. However, the abstract lacks explicit mention of real-world datasets or benchmarks, and the evaluation details are vague, making it difficult to assess generalizability.

Read-first score

Read-first score 60.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Topical relevance 42%
68.6

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=experiment,metric

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=github

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 280.

Keyword Scores

world model
9
interactive world model
9
video world model
9
generative world model
8
world dynamics prediction
7
world simulator
5
model-based reinforcement learning world model
1

Deep Analysis

Innovations

  • Explicit incorporation of geometric predictions from previous timesteps into current predictions conditioned on actions in a 4D autoregressive world model
  • Geometry-aware memory retrieval for preserving long-term spatial consistency

Methodology

DeepVerse is a 4D interactive world model that autoregressively generates video frames while explicitly incorporating geometric predictions from previous timesteps into current predictions, conditioned on actions. It uses geometric constraints to capture spatio-temporal relationships and physical dynamics, enabling geometry-aware memory retrieval for long-term spatial consistency.

Key Results

DeepVerse significantly reduces drift and enhances temporal consistency, enabling reliable generation of extended future sequences with improvements in prediction accuracy, visual realism, and scene rationality across diverse scenarios.

Tags