Mode Seeking meets Mean Seeking for Fast Long Video Generation
TLDR
Proposes a training paradigm combining mode-seeking and mean-seeking losses with a Decoupled Diffusion Transformer to generate minute-scale videos with local fidelity and long-range coherence.
Reasoning
Strengths: novel decoupling of local fidelity and long-term coherence, using limited long videos plus a frozen short-video teacher, and enabling few-step fast generation. Weaknesses: the abstract lacks quantitative details and explicit benchmark names, so empirical claims are not fully verifiable from the abstract alone.
Read-first score
Read-first score 20.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 2.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 42.