DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving
TLDR
DrivingGen is the first comprehensive benchmark for generative driving world models, with new metrics and diverse data to evaluate visual realism, trajectory plausibility, temporal coherence, and controllability.
Reasoning
The paper addresses critical gaps in evaluating driving world models by introducing a diverse dataset and novel metrics beyond generic video quality. Its strength lies in systematic benchmarking of 14 models, but the abstract cuts off before detailing results or limitations, and the benchmark's real-world impact depends on the completeness of the evaluation suite.
Read-first score
Read-first score 73.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 57.
Field roles
Rank sensitivity
Stability: volatile; rank range: 52.
Keyword Scores
Deep Analysis
Innovations
- First comprehensive benchmark for generative driving world models
- Diverse evaluation dataset curated from driving datasets and internet-scale video sources spanning varied weather, time of day, geographic regions, and complex maneuvers
- New suite of metrics jointly assessing visual realism, trajectory plausibility, temporal coherence, and controllability
- Addresses gaps in existing evaluations: generic video metrics overlook safety-critical factors, trajectory plausibility rarely quantified, temporal/agent-level consistency neglected, controllability ignored
Methodology
DrivingGen combines a diverse evaluation dataset curated from both driving datasets and internet-scale video sources, covering varied weather, time of day, geographic regions, and complex maneuvers, with a new suite of metrics that jointly assess visual realism, trajectory plausibility, temporal coherence, and controllability. The benchmark evaluates 14 state-of-the-art models using these metrics.
Key Results
Benchmarking 14 state-of-the-art models reveals clear trade-offs: general models look better but break physics, while driving-specific ones capture motion realistically but lag in visual quality.