Owl-1: Omni World Model for Consistent Long Video Generation
TLDR
Owl-1 proposes an Omni World Model using latent states and temporal dynamics for consistent long video generation.
Reasoning
The paper introduces a novel approach to long video generation by modeling world dynamics in latent space, addressing inconsistency issues. Strengths include a clear methodology and evaluation on standard benchmarks. Weaknesses are the lack of interactive or RL components and limited scope to video generation.
Read-first score
Read-first score 67.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 40.
Field roles
Rank sensitivity
Stability: volatile; rank range: 252.
Keyword Scores
Deep Analysis
Innovations
- Omni World Model (Owl-1) for consistent long video generation using a latent state variable to represent the world
- Modeling long-term developments in a latent space, where the latent state is decoded into video observations and updated by anticipated temporal dynamics
- Interaction between evolving dynamics and persistent state to enhance diversity and consistency in long videos
Methodology
Owl-1 represents the world with a latent state variable that can be decoded into explicit video observations. These observations serve as a basis for anticipating temporal dynamics, which in turn update the state variable. The model uses video generation models (VGMs) to film the latent state into videos, enabling iterative long video generation with coherent conditions.
Key Results
Owl-1 achieves comparable performance with state-of-the-art methods on VBench-I2V and VBench-Long benchmarks, validating its ability to generate high-quality video observations for long videos.