Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation

arXiv 2026 51.1 method

TLDR

ST-Gen4D embeds 4D spatiotemporal cognition into a world model for consistent 4D generation using graphs and diffusion.

Reasoning

The paper proposes a novel framework integrating spatiotemporal cognition with a world model for 4D generation, addressing both global and local dynamics. However, the abstract lacks details on empirical evaluation and comparison with baselines, and the world model's role is not fully clarified.

Read-first score

Read-first score 51.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 32.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
74.3

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.74348546

Methodology quality 18%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=dataset,experiment

Reproducibility 18%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,dataset

Topical relevance 29%
45.7

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FrontierBridge

Rank sensitivity

Stability: volatile; rank range: 237.

Keyword Scores

world model
9
world dynamics prediction
8
generative world model
6
video world model
5
world simulator
2
interactive world model
1
model-based reinforcement learning world model
1

Deep Analysis

Innovations

  • Spatiotemporal representation encoding various modalities into multiple representations as a feature basis
  • Spatiotemporal cognition sculpting representations into global appearance graph and local dynamic graph, fused via semantic-bridged spatiotemporal fusion to obtain a 4D cognition graph
  • Spatiotemporal reasoning using a world model to derive future state based on the 4D cognition
  • Spatiotemporal generation leveraging derived cognition as condition to guide latent diffusion for 4D Gaussian generation
  • Introduction of ST-4D datasets by aggregating public 4D datasets and a self-built subset

Methodology

ST-Gen4D is a 4D generation framework built on four key designs: spatiotemporal representation (encoding multiple modalities into multiple representations), spatiotemporal cognition (constructing global appearance and local dynamic graphs fused via semantic-bridged fusion into a 4D cognition graph), spatiotemporal reasoning (using a world model to predict future states from the cognition graph), and spatiotemporal generation (using the cognition as a condition for latent diffusion to produce 4D Gaussians). The model is trained and evaluated on the proposed ST-4D dataset, which aggregates public 4D data and a self-built subset.

Key Results

Extensive experiments demonstrate the superiority of ST-Gen4D across both 3D and 4D generation tasks, indicating improved structural rationality and topological consistency over existing methods.

Tags

4D generationspatiotemporal cognitionworld modelgenerative modelsdynamic topologyCV