HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving
TLDR
HoloDrive jointly generates camera images and LiDAR point clouds for autonomous driving using 2D-3D transforms and temporal prediction.
Reasoning
The paper introduces a novel framework for multi-modal street scene generation, addressing a gap in joint 2D-3D generation. Strengths include the use of BEV-to-Camera and Camera-to-BEV transforms and depth disambiguation. Weaknesses are that it focuses only on generation metrics and does not explicitly claim to be a world model, limiting relevance to the specified keywords.
Read-first score
Read-first score 37.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 7.
Field roles
Rank sensitivity
Stability: volatile; rank range: 178.
Keyword Scores
Deep Analysis
Innovations
- Joint 2D-3D multi-modal generation of camera images and LiDAR point clouds for autonomous driving
- BEV-to-Camera and Camera-to-BEV transform modules between heterogeneous generative models
- Depth prediction branch in the 2D generative model to disambiguate un-projecting from image space to BEV space
- Extension to future prediction via temporal structure and progressive training
Methodology
HoloDrive employs BEV-to-Camera and Camera-to-BEV transform modules to bridge heterogeneous generative models for joint generation of camera images and LiDAR point clouds. A depth prediction branch is introduced in the 2D generative model to resolve ambiguity when un-projecting from image space to BEV space. The framework is extended to future prediction by adding temporal structure and using carefully designed progressive training.
Key Results
Experiments on single frame generation and world model benchmarks demonstrate that HoloDrive achieves significant performance gains over state-of-the-art methods in terms of generation metrics.