MoVieDrive: Multi-Modal Multi-View Urban Scene Video Generation
TLDR
Proposes a unified diffusion transformer for multi-modal multi-view driving video generation, achieving compelling quality and controllability.
Reasoning
Strengths include a novel unified framework for multi-modal generation and validation on real-world data. Weaknesses are that it focuses solely on video generation without claiming world modeling, dynamics prediction, or interactivity, limiting relevance to the specified keywords.
Read-first score
Read-first score 37.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 7.
Field roles
Rank sensitivity
Stability: volatile; rank range: 256.
Keyword Scores
Deep Analysis
Innovations
- Unified diffusion transformer model for multi-modal multi-view video generation in autonomous driving
- Modal-shared and modal-specific components to handle different modalities (RGB, depth, semantic maps) within a single framework
- Diverse conditioning inputs to encode controllable scene structure and content cues into the unified model
Methodology
The approach constructs a unified diffusion transformer model composed of modal-shared and modal-specific components. It leverages diverse conditioning inputs to encode controllable scene structure and content cues, enabling multi-modal multi-view driving scene video generation in a single framework.
Key Results
Experiments on a real-world autonomous driving dataset demonstrate compelling video generation quality and controllability compared to state-of-the-art methods, while supporting multi-modal multi-view data generation.