Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation

AAAI 25 2025 48.6 method

TLDR

3D4D is an interactive 4D visualization framework that generates coherent 4D scenes from images and text using WebGL and foveated rendering.

Reasoning

The paper introduces a novel interactive 4D visualization framework with real-time rendering, but lacks explicit evaluation on real-world benchmarks or connection to world dynamics prediction and reinforcement learning. Its strengths lie in the integration of WebGL and foveated rendering for user-driven exploration, while weaknesses include limited evidence of empirical validation and absence of core world model capabilities beyond visualization.

Read-first score

Read-first score 48.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 25.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=experiment,result

Reproducibility 25%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,github

Topical relevance 42%
35.7

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 409.

Keyword Scores

interactive world model
8
video world model
7
world model
6
generative world model
4
world simulator
0
world dynamics prediction
0
model-based reinforcement learning world model
0

Deep Analysis

Innovations

  • Integration of WebGL with Supersplat rendering for 4D visualization
  • Four core modules to transform static images and text into coherent 4D scenes
  • Foveated rendering strategy for efficient real-time multi-modal interaction
  • Adaptive, user-driven exploration of complex 4D environments

Methodology

The framework integrates WebGL with Supersplat rendering and employs a foveated rendering strategy. It transforms static images and text into coherent 4D scenes through four core modules, enabling real-time multi-modal interaction.

Key Results

No experimental results are reported in the abstract; the framework's capabilities are described qualitatively.

Tags