DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT
TLDR
DrivingWorld is a GPT-style world model for autonomous driving that generates long-duration, high-fidelity video sequences using spatial-temporal fusion.
Reasoning
Strengths include novel spatial-temporal fusion mechanisms and achieving longer video generation than prior work. Weaknesses are limited quantitative results and lack of detail on evaluation metrics beyond duration.
Read-first score
Read-first score 68.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 45.
Field roles
Rank sensitivity
Stability: volatile; rank range: 146.
Keyword Scores
Deep Analysis
Innovations
- Spatial-temporal fusion mechanisms for GPT-style world model
- Next-state prediction strategy to model temporal coherence between consecutive frames
- Next-token prediction strategy to capture spatial information within each frame
- Novel masking strategy and reweighting strategy for token prediction to mitigate long-term drifting and enable precise control
Methodology
DrivingWorld is a GPT-style world model for autonomous driving that integrates spatial-temporal fusion mechanisms. It uses a next-state prediction strategy to model temporal coherence between consecutive frames and a next-token prediction strategy to capture spatial information within each frame. Additionally, a novel masking strategy and reweighting strategy are applied to token prediction to mitigate long-term drifting and enable precise control.
Key Results
The model generates high-fidelity and consistent video clips of over 40 seconds in duration, which is over 2 times longer than state-of-the-art driving world models, achieving superior visual quality and significantly more accurate controllable future video generation.