OmniGen: Unified Multimodal Sensor Generation for Autonomous Driving
TLDR
OmniGen generates aligned multimodal sensor data for autonomous driving using a shared BEV space and volume rendering.
Reasoning
The paper proposes a unified framework for multimodal sensor generation with a novel reconstruction method, but it is limited to autonomous driving and lacks explicit real-world evaluation in the abstract.
Read-first score
Read-first score 31, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 3.
Field roles
Rank sensitivity
Stability: volatile; rank range: 104.
Keyword Scores
Deep Analysis
Innovations
- Unified framework for generating aligned multimodal sensor data (LiDAR and multi-view cameras) using a shared Bird's Eye View (BEV) space.
- Novel generalizable multimodal reconstruction method (UAE) that jointly decodes LiDAR and multi-view camera data via volume rendering.
- Incorporation of a Diffusion Transformer (DiT) with a ControlNet branch for controllable multimodal sensor generation.
Methodology
OmniGen uses a shared Bird's Eye View (BEV) space to unify multimodal features from LiDAR and multi-view cameras. It designs a novel multimodal reconstruction method called UAE, which employs volume rendering to jointly decode the sensor data. A Diffusion Transformer (DiT) with a ControlNet branch is integrated to enable controllable generation of multimodal sensor data.
Key Results
Comprehensive experiments demonstrate that OmniGen achieves desired performances in unified multimodal sensor data generation, with multimodal consistency and flexible sensor adjustments.