LongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoE
TLDR
LongScape combines intra-chunk diffusion and inter-chunk autoregression with action-guided chunking and Context-aware MoE for stable long-horizon video generation in embodied world models.
Reasoning
The paper presents a novel hybrid framework that addresses temporal inconsistency and visual drift in long-horizon video generation, with a clear methodological contribution (action-guided chunking and CMoE). However, the abstract lacks explicit mention of real-world benchmarks or datasets, and the evaluation details are not provided, limiting assessment of empirical rigor.
Read-first score
Read-first score 74, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 53.
Field roles
Rank sensitivity
Stability: volatile; rank range: 48.
Keyword Scores
Deep Analysis
Innovations
- Hybrid framework combining intra-chunk diffusion denoising with inter-chunk autoregressive causal generation
- Action-guided, variable-length chunking mechanism based on semantic context of robotic actions
- Context-aware Mixture-of-Experts (CMoE) that adaptively activates specialized experts per chunk
Methodology
LongScape is a hybrid framework that adaptively combines intra-chunk diffusion denoising with inter-chunk autoregressive causal generation. It uses an action-guided, variable-length chunking mechanism that partitions video based on semantic context of robotic actions, and a Context-aware Mixture-of-Experts (CMoE) that adaptively activates specialized experts for each chunk during generation.
Key Results
Extensive experimental results demonstrate that the method achieves stable and consistent long-horizon generation over extended rollouts.