Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control

arXiv 25.3 2025 32.1 method, application

TLDR

MagicDrive-V2 generates high-resolution long multi-view driving videos with geometric and textual control using MVDiT and progressive training.

Reasoning

The paper introduces novel techniques for controllable driving video generation, but does not explicitly frame itself as a world model or address dynamics prediction or interactive simulation. Strengths include multi-view and geometric control; weaknesses include lack of direct relevance to world model terminology.

Read-first score

Read-first score 32.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 4.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=experiment,metric

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Topical relevance 42%
5.7

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 123.

Keyword Scores

video world model
2
world model
1
generative world model
1
world simulator
0
interactive world model
0
world dynamics prediction
0
model-based reinforcement learning world model
0

Deep Analysis

Innovations

  • Integration of MVDiT block and spatial-temporal conditional encoding to enable multi-view video generation and precise geometric control
  • Efficient method for obtaining contextual descriptions for videos to support diverse textual control
  • Progressive training strategy using mixed video data to enhance training efficiency and generalizability

Methodology

MagicDrive-V2 builds on the DiT with 3D VAE framework and introduces MVDiT blocks and spatial-temporal conditional encoding to achieve multi-view video generation and precise geometric control. It also employs an efficient method to extract contextual descriptions from videos for textual control, and uses a progressive training strategy with mixed video data to improve efficiency and generalization.

Key Results

MagicDrive-V2 enables multi-view driving video synthesis with 3.3× resolution and 4× frame count compared to current state-of-the-art, while supporting rich contextual and geometric controls.

Tags