Awesome World Model Hub 论文 · 数据集 · 项目
← 返回论文列表

DrivingGen:自动驾驶中生成式视频世界模型的综合基准

arXiv 26.1 2026 73.7 benchmark, application

TLDR

DrivingGen是首个生成式驾驶世界模型综合基准,用新指标和多样化数据评估视觉、轨迹、时间一致性和可控性。

评分理由

The paper addresses critical gaps in evaluating driving world models by introducing a diverse dataset and novel metrics beyond generic video quality. Its strength lies in systematic benchmarking of 14 models, but the abstract cuts off before detailing results or limitations, and the benchmark's real-world impact depends on the completeness of the evaluation suite.

Read-first 评分解释

综合优先阅读分 73.7,由主题、引用、图谱、方法、可复现性和近期性等信号加权得到。 原始总分保留为 57。

近期性 8%
100

使用温和的时间衰减,让近期论文更容易浮现,同时保留较早基础工作的价值。 年份:2026

主题相关性 42%
81.4

使用现有 LLM 关键词相关性评分,并归一化到 0-100。 关键词:world model、world simulator、generative world model、interactive world model、video world model、world dynamics prediction、model-based reinforcement learning world model

方法质量 25%
80

检查可见的摘要与分析字段,寻找实验、数据集、基线、指标和局限性等方法证据。 命中信号:基准、数据集、评估、指标

可复现性 25%
46

检查链接和可见文本中的论文、代码、数据集、工件与仓库信号。 论文:有;代码:无;数据:无;命中信号:数据集、GitHub

研究版图角色

前沿论文方法锚点

排序敏感性

稳定性:volatile;排名波动范围:52。

关键词评分

world model
10
generative world model
10
video world model
10
world simulator
8
world dynamics prediction
8
interactive world model
7
model-based reinforcement learning world model
4

深度分析

创新点

  • 首个针对生成式驾驶世界模型的综合基准
  • 从驾驶数据集和互联网规模视频源中精选的多样化评估数据集,涵盖多种天气、时段、地理区域和复杂操作
  • 新的一套指标,联合评估视觉真实感、轨迹合理性、时间一致性和可控性
  • 解决了现有评估中的空白:通用视频指标忽视安全关键因素,轨迹合理性很少量化,时间/代理级别一致性被忽略,可控性被忽视

方法

DrivingGen结合了从驾驶数据集和互联网规模视频源中精选的多样化评估数据集,涵盖多种天气、时段、地理区域和复杂操作,以及一套新指标,联合评估视觉真实感、轨迹合理性、时间一致性和可控性。该基准使用这些指标评估了14个最先进模型。

关键结果

对14个最先进模型的基准测试揭示了明显的权衡:通用模型外观更好但违反物理规律,而驾驶专用模型能真实捕捉运动但在视觉质量上落后。

标签