MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
TLDR
Introduces Mixture-of-Groups Attention, a learnable sparse attention method for efficient long video generation, enabling minute-level 480p videos with long context.
Reasoning
The paper proposes a novel sparse attention mechanism with a learnable token router, addressing quadratic attention costs in long video generation and showing strong empirical results. However, the abstract provides limited detail on baseline comparisons or limitations, and the work is not framed as a world model.
Read-first score
Read-first score 24, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 3.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 76.