ST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D Generation
TLDR
ST-Gen4D embeds 4D spatiotemporal cognition into a world model for consistent 4D generation using graphs and diffusion.
Reasoning
The paper proposes a novel framework integrating spatiotemporal cognition with a world model for 4D generation, addressing both global and local dynamics. However, the abstract lacks details on empirical evaluation and comparison with baselines, and the world model's role is not fully clarified.
Read-first score
Read-first score 51.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 32.
Field roles
Rank sensitivity
Stability: volatile; rank range: 237.
Keyword Scores
Deep Analysis
Innovations
- Spatiotemporal representation encoding various modalities into multiple representations as a feature basis
- Spatiotemporal cognition sculpting representations into global appearance graph and local dynamic graph, fused via semantic-bridged spatiotemporal fusion to obtain a 4D cognition graph
- Spatiotemporal reasoning using a world model to derive future state based on the 4D cognition
- Spatiotemporal generation leveraging derived cognition as condition to guide latent diffusion for 4D Gaussian generation
- Introduction of ST-4D datasets by aggregating public 4D datasets and a self-built subset
Methodology
ST-Gen4D is a 4D generation framework built on four key designs: spatiotemporal representation (encoding multiple modalities into multiple representations), spatiotemporal cognition (constructing global appearance and local dynamic graphs fused via semantic-bridged fusion into a 4D cognition graph), spatiotemporal reasoning (using a world model to predict future states from the cognition graph), and spatiotemporal generation (using the cognition as a condition for latent diffusion to produce 4D Gaussians). The model is trained and evaluated on the proposed ST-4D dataset, which aggregates public 4D data and a self-built subset.
Key Results
Extensive experiments demonstrate the superiority of ST-Gen4D across both 3D and 4D generation tasks, indicating improved structural rationality and topological consistency over existing methods.