Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks
TLDR
Dream4Drive uses a driving world model to generate synthetic multi-view videos, enhancing downstream perception tasks and corner case detection in autonomous driving.
Reasoning
The paper addresses a gap in evaluating driving world models for downstream perception, introducing a flexible framework for generating corner cases. Strengths include a novel 3D-aware generation pipeline and a new dataset, but the abstract lacks quantitative results and comparisons to baselines.
Read-first score
Read-first score 67.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 40.
Field roles
Rank sensitivity
Stability: volatile; rank range: 213.
Keyword Scores
Deep Analysis
Innovations
- Rethinking driving world models as synthetic data generators for downstream perception tasks, not just generation quality and controllability.
- Dream4Drive framework that decomposes input video into 3D-aware guidance maps, renders 3D assets, and fine-tunes world model to produce edited multi-view photorealistic videos.
- Ability to generate multi-view corner cases at scale for autonomous driving perception.
- Contribution of DriveObj3D, a large-scale 3D asset dataset covering typical driving categories for diverse 3D-aware video editing.
Methodology
Dream4Drive first decomposes the input video into several 3D-aware guidance maps, then renders 3D assets onto these guidance maps. Finally, the driving world model is fine-tuned to produce edited, multi-view photorealistic videos, which are used to train downstream perception models. The framework is evaluated through comprehensive experiments on downstream perception tasks under various training epochs.
Key Results
Dream4Drive effectively boosts the performance of downstream perception models under various training epochs, particularly for corner case perception in autonomous driving.