Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
TLDR
Introduces Hybrid Memory and HyDRA for video world models to maintain subject consistency when hidden, with a new dataset HM-World.
Reasoning
The paper addresses a clear gap in video world models by proposing a hybrid memory paradigm and a specialized architecture (HyDRA) that outperforms existing methods on a new large-scale dataset (HM-World). However, the contribution is limited to video generation tasks and does not extend to interactive or reinforcement learning settings, and the term 'world model' may be overstated for a video prediction context.
Read-first score
Read-first score 69.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 42.
Field roles
Rank sensitivity
Stability: volatile; rank range: 220.
Keyword Scores
Deep Analysis
Innovations
- Hybrid Memory paradigm requiring models to act as precise archivists for static backgrounds and vigilant trackers for dynamic subjects, ensuring motion continuity during out-of-view intervals.
- HM-World, the first large-scale video dataset dedicated to hybrid memory, featuring 59K high-fidelity clips with decoupled camera and subject trajectories, 17 scenes, 49 subjects, and meticulously designed exit-entry events.
- HyDRA, a specialized memory architecture that compresses memory into tokens and utilizes a spatiotemporal relevance-driven retrieval mechanism to preserve identity and motion of hidden subjects.
Methodology
The paper proposes HyDRA, a memory architecture that compresses memory into tokens and employs a spatiotemporal relevance-driven retrieval mechanism to selectively attend to relevant motion cues. The method is evaluated on the newly constructed HM-World dataset, which includes 59K clips with decoupled camera and subject trajectories and exit-entry events. The approach is compared against state-of-the-art methods in dynamic video world modeling.
Key Results
Extensive experiments on HM-World demonstrate that HyDRA significantly outperforms state-of-the-art approaches in both dynamic subject consistency and overall generation quality.