DriveGenVLM: Real-world Video Generation for Vision Language Model based Autonomous Driving
TLDR
DriveGenVLM generates driving videos using diffusion models and uses VLMs for scene understanding to enhance autonomous driving.
Reasoning
The paper integrates video generation with VLMs, using the Waymo dataset and FVD evaluation, which are strengths. However, it lacks novelty and does not explicitly address world models or interactive aspects, limiting its scope.
Read-first score
Read-first score 32.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 3.
Field roles
Rank sensitivity
Stability: volatile; rank range: 69.
Keyword Scores
Deep Analysis
Innovations
- Integration of video generation with vision language models (VLMs) for autonomous driving
- Use of denoising diffusion probabilistic models (DDPM) to generate real-world driving video sequences
- Application of pre-trained EILEV model to provide narrations for generated driving videos
Methodology
The DriveGenVLM framework employs a denoising diffusion probabilistic model (DDPM) for video generation, trained on the Waymo open dataset. Generated videos are evaluated using the Fréchet Video Distance (FVD) score. A pre-trained model, Efficient In-context Learning on Egocentric Videos (EILEV), is then used to produce corresponding narrations for the generated videos.
Key Results
The paper reports that the diffusion model is evaluated using the FVD score to ensure quality and realism, but no specific numerical results are provided in the abstract. The narrations from EILEV are claimed to enhance traffic scene understanding, navigation, and planning.