Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving

arXiv 26.1 2026 73.7 benchmark, application

TLDR

DrivingGen is the first comprehensive benchmark for generative driving world models, with new metrics and diverse data to evaluate visual realism, trajectory plausibility, temporal coherence, and controllability.

Reasoning

The paper addresses critical gaps in evaluating driving world models by introducing a diverse dataset and novel metrics beyond generic video quality. Its strength lies in systematic benchmarking of 14 models, but the abstract cuts off before detailing results or limitations, and the benchmark's real-world impact depends on the completeness of the evaluation suite.

Read-first score

Read-first score 73.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 57.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Topical relevance 42%
81.4

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,dataset,evaluation,metric

Reproducibility 25%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset,github

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 52.

Keyword Scores

world model
10
generative world model
10
video world model
10
world simulator
8
world dynamics prediction
8
interactive world model
7
model-based reinforcement learning world model
4

Deep Analysis

Innovations

  • First comprehensive benchmark for generative driving world models
  • Diverse evaluation dataset curated from driving datasets and internet-scale video sources spanning varied weather, time of day, geographic regions, and complex maneuvers
  • New suite of metrics jointly assessing visual realism, trajectory plausibility, temporal coherence, and controllability
  • Addresses gaps in existing evaluations: generic video metrics overlook safety-critical factors, trajectory plausibility rarely quantified, temporal/agent-level consistency neglected, controllability ignored

Methodology

DrivingGen combines a diverse evaluation dataset curated from both driving datasets and internet-scale video sources, covering varied weather, time of day, geographic regions, and complex maneuvers, with a new suite of metrics that jointly assess visual realism, trajectory plausibility, temporal coherence, and controllability. The benchmark evaluates 14 state-of-the-art models using these metrics.

Key Results

Benchmarking 14 state-of-the-art models reveals clear trade-offs: general models look better but break physics, while driving-specific ones capture motion realistically but lag in visual quality.

Tags