Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment

ACM MM 25 2025 52.2 method

TLDR

Collecting multi-view driving scenario videos to enhance the performance of 3D visual perception tasks presents significant challenges and incurs substantial costs, making generative models for realistic data an appealing alternative.

Reasoning

Fallback reasoning generated from available title and abstract metadata: Collecting multi-view driving scenario videos to enhance the performance of 3D visual perception tasks presents significant challenges and incurs substantial costs, making generative models for realistic data an appealing alternative. Yet, the videos generated by recent works suffer...

Read-first score

Read-first score 52.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Reproducibility 25%
81

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=dataset,github

Methodology quality 25%
50

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=dataset

Topical relevance 42%
29.4

Matches configured research keywords against title, abstract, tags, and analysis text. matched=6

Field roles

FrontierReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 471.

Deep Analysis

Innovations

  • Multi-Control Auxiliary Branch Distillation
  • Resolution Progressive Sampling

Methodology

DiVE is a diffusion transformer-based generative framework that uses unified cross-attention and a SketchFormer for precise multimodal control over bird's-eye view layouts and textual descriptions, and incorporates a view-inflated attention mechanism for cross-view consistency without adding extra parameters. It also introduces Multi-Control Auxiliary Branch Distillation to streamline multi-condition classifier-free guidance and Resolution Progressive Sampling, a training-free acceleration strategy that staggers resolution scaling to reduce latency.

Key Results

On the nuScenes dataset, DiVE achieves state-of-the-art performance in multi-view video generation, producing photorealistic outputs with exceptional temporal and cross-view coherence, while achieving a 2.62x speedup with minimal quality degradation.

Tags