Reuse and Diffuse: Iterative Denoising for Text-to-Video Generation
TLDR
Proposes VidRD, a framework using latent diffusion models to iteratively generate additional video frames for text-to-video generation with improved temporal consistency.
Reasoning
The paper presents a clear method for extending video frames using latent diffusion and includes quantitative and qualitative evaluations. However, the abstract lacks explicit discussion of limitations or comparisons to broader world model frameworks, and the connection to world modeling is indirect.
Read-first score
Read-first score 35.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 5.
Field roles
Candidate
Rank sensitivity
Stability: volatile; rank range: 90.