Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation

arXiv 2025 24 method

TLDR

Introduces Mixture-of-Groups Attention, a learnable sparse attention method for efficient long video generation, enabling minute-level 480p videos with long context.

Reasoning

The paper proposes a novel sparse attention mechanism with a learnable token router, addressing quadratic attention costs in long video generation and showing strong empirical results. However, the abstract provides limited detail on baseline comparisons or limitations, and the work is not framed as a world model.

Read-first score

Read-first score 24, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 3.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
30

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=experiment

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Topical relevance 42%
4.3

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 76.

Keyword Scores

video world model
2
world model
1
world simulator
0
generative world model
0
interactive world model
0
world dynamics prediction
0
model-based reinforcement learning world model
0

Tags