LIVE: Long-horizon Interactive Video World Modeling
TLDR
LIVE introduces a cycle-consistency objective for long-horizon video world models, eliminating teacher-based distillation and achieving state-of-the-art performance.
Reasoning
The paper presents a novel cycle-consistency approach to bound error accumulation in autoregressive video world models, which is a clear strength. However, the abstract lacks explicit details on real-world datasets and does not discuss limitations or comparisons with baselines in depth.
Read-first score
Read-first score 66.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 60.
Field roles
Rank sensitivity
Stability: volatile; rank range: 331.
Keyword Scores
Deep Analysis
Innovations
- Cycle-consistency objective that enforces bounded error accumulation without teacher-based distillation
- Unified view encompassing different approaches to long-horizon video world modeling
- Progressive training curriculum to stabilize training
Methodology
LIVE performs a forward rollout from ground-truth frames, then applies a reverse generation process to reconstruct the initial state. The diffusion loss is computed on the reconstructed terminal state, providing an explicit constraint on long-horizon error propagation.
Key Results
LIVE achieves state-of-the-art performance on long-horizon benchmarks, generating stable, high-quality videos far beyond training rollout lengths.