Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving

CVPR 24 2024 76.1 method, application

TLDR

Drive-WM: a driving world model that generates multiview videos for safe planning by forecasting multiple futures and selecting optimal trajectories via image-based rewards.

Reasoning

The paper introduces a novel world model for autonomous driving that generates high-fidelity multiview videos and enables planning by evaluating multiple future scenarios. Strengths include real-world dataset evaluation and compatibility with end-to-end planning models; weaknesses are limited explicit discussion of interactive or RL-based components.

Read-first score

Read-first score 76.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 58.

Topical relevance 42%
82.9

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 25%
81

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=dataset,github

Recency 8%
75.1

Uses a gentle age decay so recent papers surface without erasing older foundations. 2024

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=dataset,evaluation

Field roles

Reproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 57.

Keyword Scores

world model
10
generative world model
9
video world model
9
world dynamics prediction
9
interactive world model
8
world simulator
7
model-based reinforcement learning world model
6

Deep Analysis

Innovations

  • First driving world model compatible with existing end-to-end planning models
  • Joint spatial-temporal modeling via view factorization for high-fidelity multiview video generation
  • Application of world model for safe driving planning by generating multiple futures and selecting optimal trajectory based on image-based rewards

Methodology

Drive-WM employs joint spatial-temporal modeling facilitated by view factorization to generate high-fidelity multiview videos of driving scenes. It enables driving into multiple futures based on distinct driving maneuvers and determines the optimal trajectory according to image-based rewards. The model is evaluated on real-world driving datasets.

Key Results

The method generates high-quality, consistent, and controllable multiview videos, demonstrating potential for real-world simulations and safe planning.

Tags