MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving
TLDR
MiLA generates high-fidelity, long-duration driving videos up to one minute using coarse-to-refine and denoising modules, achieving SOTA on nuScenes.
Reasoning
The paper proposes a novel framework for long-term video generation in autonomous driving, addressing error accumulation with coarse-to-refine and denoising modules. Strengths include state-of-the-art results on nuScenes, but weaknesses are limited evaluation to a single dataset and lack of real-world deployment or interactive capabilities.
Read-first score
Read-first score 59.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 43.
Field roles
Rank sensitivity
Stability: volatile; rank range: 214.
Keyword Scores
Deep Analysis
Innovations
- Coarse-to-Re(fine) approach for stabilizing video generation and correcting distortion of dynamic objects
- Temporal Progressive Denoising Scheduler
- Joint Denoising and Correcting Flow modules
Methodology
MiLA is a framework for long-term video generation using a Coarse-to-Re(fine) approach to stabilize generation and correct dynamic object distortion. It incorporates a Temporal Progressive Denoising Scheduler and Joint Denoising and Correcting Flow modules. The model is trained and evaluated on the nuScenes dataset, achieving state-of-the-art performance.
Key Results
MiLA achieves state-of-the-art performance in video generation quality on the nuScenes dataset.