DriveLaW: Unifying Planning and Video Generation in a Latent Driving World
TLDR
DriveLaW unifies video generation and motion planning for autonomous driving by injecting latent representations from a world model into a diffusion planner.
Reasoning
The paper presents a novel unified paradigm with strong empirical results on video prediction and planning benchmarks. However, the abstract lacks details on limitations and the real-world scope of the benchmarks.
Read-first score
Read-first score 49.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.
Field roles
Rank sensitivity
Stability: volatile; rank range: 438.
Keyword Scores
Deep Analysis
Innovations
- Unifying video generation and motion planning by directly injecting latent representation from video generator into planner
- DriveLaW-Video: a powerful world model that generates high-fidelity forecasting with expressive latent representations
- DriveLaW-Act: a diffusion planner that generates consistent and reliable trajectories from the latent of DriveLaW-Video
- Three-stage progressive training strategy to optimize both components
Methodology
DriveLaW consists of two core components: DriveLaW-Video, a world model for high-fidelity video forecasting with expressive latent representations, and DriveLaW-Act, a diffusion planner that generates trajectories from the latent of DriveLaW-Video. Both components are optimized using a three-stage progressive training strategy. The model unifies planning and video generation by directly injecting the latent representation from the video generator into the planner.
Key Results
DriveLaW achieves state-of-the-art results on video prediction, surpassing the best-performing work by 33.3% in FID and 1.8% in FVD, and also sets a new record on the NAVSIM planning benchmark.