Beyond Pixel Histories: World Models with Persistent 3D State
TLDR
PERSIST introduces a world model with persistent 3D state, improving spatial memory and 3D consistency for interactive video generation.
Reasoning
The paper presents a novel paradigm that explicitly maintains a latent 3D scene, addressing key limitations of prior interactive world models. Strengths include clear methodology, quantitative and qualitative evaluations, and novel capabilities like single-image 3D synthesis. Weaknesses are not evident from the abstract alone, but the approach may have computational overhead or scalability concerns not discussed.
Read-first score
Read-first score 52.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 56.
Field roles
Rank sensitivity
Stability: volatile; rank range: 484.
Keyword Scores
Deep Analysis
Innovations
- Introduces PERSIST, a world model paradigm that simulates the evolution of a latent 3D scene (environment, camera, and renderer) for persistent spatial memory and consistent geometry.
- Enables synthesizing diverse 3D environments from a single image.
- Supports fine-grained, geometry-aware control over generated experiences via environment editing and specification directly in 3D space.
Methodology
PERSIST models the world by simulating the evolution of a latent 3D scene comprising environment, camera, and renderer. This allows synthesis of new frames with persistent spatial memory and consistent geometry, moving beyond pixel-level histories to a 3D state representation.
Key Results
Quantitative metrics and a qualitative user study demonstrate substantial improvements in spatial memory, 3D consistency, and long-horizon stability over existing methods, enabling coherent, evolving 3D worlds.