Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

MUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric Representations

arXiv 23.11 2023 54.5 method

TLDR

MUVO combines multimodal sensor data (camera, lidar) with 3D occupancy prediction for a generative world model in autonomous driving.

Reasoning

The paper addresses a gap by integrating lidar and camera data with 3D occupancy prediction, but the abstract lacks details on real-world validation and does not mention interactive or RL aspects.

Read-first score

Read-first score 54.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 32.

Methodology quality 25%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,evaluation,experiment,metric,result

Recency 8%
65.1

Uses a gentle age decay so recent papers surface without erasing older foundations. 2023

Topical relevance 42%
45.7

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

Methodology anchor

Rank sensitivity

Stability: volatile; rank range: 222.

Keyword Scores

world model
10
generative world model
10
world dynamics prediction
7
world simulator
3
video world model
2
interactive world model
0
model-based reinforcement learning world model
0

Deep Analysis

Innovations

  • Combining multimodal sensor data (camera and lidar) with 3D occupancy prediction in a generative world model for autonomous driving.
  • Systematic evaluation of different sensor fusion strategies within a world model framework.
  • Analysis of weaknesses in current sensor fusion approaches and demonstration of benefits from additionally predicting 3D occupancy.

Methodology

MUVO is a multimodal generative world model that uses geometric voxel representations. It integrates camera and lidar data and predicts both raw sensor outputs and 3D occupancy. The experiments compare various sensor fusion strategies to assess their impact on prediction quality and to identify limitations of existing fusion methods.

Key Results

The abstract does not provide specific quantitative results; it describes experimental evaluations that examine the effects of sensor fusion strategies and the advantages of incorporating 3D occupancy prediction, but exact metrics are not stated.

Tags