UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving
TLDR
UniDrive-WM is a unified VLM-based world model for autonomous driving that jointly performs scene understanding, trajectory planning, and future image generation.
Reasoning
The paper's strength lies in tightly integrating perception, planning, and generation within a single VLM architecture, achieving significant improvements on the Bench2Drive benchmark. However, it is limited to a single benchmark and lacks real-world deployment or diverse scenario validation.
Read-first score
Read-first score 42.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 42.
Field roles
Rank sensitivity
Stability: volatile; rank range: 329.
Keyword Scores
Deep Analysis
Innovations
- Unified VLM-based world model that jointly performs driving-scene understanding, trajectory planning, and trajectory-conditioned future image generation within a single architecture
- Trajectory planner predicts future trajectory which conditions a VLM-based image generator to produce plausible future frames, providing additional supervisory signals to enhance scene understanding and iteratively refine trajectory generation
- Comparison of discrete and continuous output representations for future image prediction and their influence on downstream driving performance
Methodology
UniDrive-WM is a unified VLM-based world model that jointly performs scene understanding, trajectory planning, and trajectory-conditioned future image generation. The trajectory planner predicts a future trajectory, which conditions a VLM-based image generator to produce future frames; these predictions provide supervisory signals to enhance understanding and iteratively refine trajectory generation. The model is evaluated on the Bench2Drive benchmark.
Key Results
UniDrive-WM improves planning performance by 7.3% in L2 trajectory error and 10.4% in collision rate over the previous best method on the Bench2Drive benchmark, while producing high-fidelity future images.