Generating Long Videos of Dynamic Scenes
TLDR
A video generation model that redesigns temporal latent representation and uses two-phase training to generate long videos with consistent object motion, camera viewpoint, and new content.
Reasoning
The paper addresses a key limitation in video generation—long-term temporal consistency—by introducing a redesigned temporal latent representation and a two-phase training strategy, and it contributes new benchmark datasets. However, the abstract does not report quantitative results or comparisons, and the proposed benchmarks may not fully capture real-world complexity.
Read-first score
Read-first score 32.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 9.
Field roles
Candidate
Rank sensitivity
Stability: volatile; rank range: 52.