Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

MoVieDrive: Multi-Modal Multi-View Urban Scene Video Generation

arXiv 25.8 2025 37.9 method, application

TLDR

Proposes a unified diffusion transformer for multi-modal multi-view driving video generation, achieving compelling quality and controllability.

Reasoning

Strengths include a novel unified framework for multi-modal generation and validation on real-world data. Weaknesses are that it focuses solely on video generation without claiming world modeling, dynamics prediction, or interactivity, limiting relevance to the specified keywords.

Read-first score

Read-first score 37.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 7.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=dataset,experiment

Reproducibility 25%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,dataset

Topical relevance 42%
10

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 256.

Keyword Scores

video world model
3
generative world model
2
world model
1
world simulator
1
interactive world model
0
world dynamics prediction
0
model-based reinforcement learning world model
0

Deep Analysis

Innovations

  • Unified diffusion transformer model for multi-modal multi-view video generation in autonomous driving
  • Modal-shared and modal-specific components to handle different modalities (RGB, depth, semantic maps) within a single framework
  • Diverse conditioning inputs to encode controllable scene structure and content cues into the unified model

Methodology

The approach constructs a unified diffusion transformer model composed of modal-shared and modal-specific components. It leverages diverse conditioning inputs to encode controllable scene structure and content cues, enabling multi-modal multi-view driving scene video generation in a single framework.

Key Results

Experiments on a real-world autonomous driving dataset demonstrate compelling video generation quality and controllability compared to state-of-the-art methods, while supporting multi-modal multi-view data generation.

Tags