SimGen: Simulator-conditioned Driving Scene Generation
TLDR
Controllable synthetic data generation can substantially lower the annotation cost of training data.
Reasoning
Fallback reasoning generated from available title and abstract metadata: Controllable synthetic data generation can substantially lower the annotation cost of training data. Prior works use diffusion models to generate driving images conditioned on the 3D object layout. However, those models are trained on small-scale datasets like nuScenes,...
Read-first score
Read-first score 53.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals.
Field roles
Rank sensitivity
Stability: volatile; rank range: 488.
Deep Analysis
Innovations
- Simulator-conditioned scene generation framework that mixes data from simulator and real world
- Novel cascade diffusion pipeline to address sim-to-real gaps and multi-condition conflicts
- DIVA dataset: 147.5 hours of real-world driving videos from 73 locations worldwide and simulated data from MetaDrive
Methodology
SimGen uses a cascade diffusion pipeline conditioned on 3D object layout and text prompts, trained on a mixture of real-world driving videos (DIVA) and simulated data from MetaDrive. The cascade design handles the sim-to-real gap and resolves conflicts between multiple conditioning signals, enabling controllable generation of diverse driving scenes.
Key Results
SimGen achieves superior generation quality and diversity while preserving controllability, and improves performance on BEV detection and segmentation tasks through synthetic data augmentation. It also demonstrates capability in generating safety-critical driving data.