DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment
TLDR
Collecting multi-view driving scenario videos to enhance the performance of 3D visual perception tasks presents significant challenges and incurs substantial costs, making generative models for realistic data an appealing alternative.
Reasoning
Fallback reasoning generated from available title and abstract metadata: Collecting multi-view driving scenario videos to enhance the performance of 3D visual perception tasks presents significant challenges and incurs substantial costs, making generative models for realistic data an appealing alternative. Yet, the videos generated by recent works suffer...
Read-first score
Read-first score 52.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals.
Field roles
Rank sensitivity
Stability: volatile; rank range: 471.
Deep Analysis
Innovations
- Multi-Control Auxiliary Branch Distillation
- Resolution Progressive Sampling
Methodology
DiVE is a diffusion transformer-based generative framework that uses unified cross-attention and a SketchFormer for precise multimodal control over bird's-eye view layouts and textual descriptions, and incorporates a view-inflated attention mechanism for cross-view consistency without adding extra parameters. It also introduces Multi-Control Auxiliary Branch Distillation to streamline multi-condition classifier-free guidance and Resolution Progressive Sampling, a training-free acceleration strategy that staggers resolution scaling to reduce latency.
Key Results
On the nuScenes dataset, DiVE achieves state-of-the-art performance in multi-view video generation, producing photorealistic outputs with exceptional temporal and cross-view coherence, while achieving a 2.62x speedup with minimal quality degradation.