Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Probing Multimodal LLMs as World Models for Driving

arXiv 24.5 2024 60.8 benchmark, system, application

TLDR

MLLMs like GPT-4o struggle to understand dynamic driving scenes across frames, despite excelling at single images, highlighting gaps in world model capabilities.

Reasoning

Strengths: challenges common assumptions with systematic evaluation using a new dataset and simulator. Weaknesses: only identifies limitations without proposing solutions, and evaluation is limited to specific MLLMs.

Read-first score

Read-first score 60.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 24.

Reproducibility 25%
81

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=dataset,github

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=dataset,evaluation,experiment

Recency 8%
75.1

Uses a gentle age decay so recent papers surface without erasing older foundations. 2024

Topical relevance 42%
34.3

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Field roles

Methodology anchorReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 645.

Keyword Scores

world model
9
world dynamics prediction
7
world simulator
4
video world model
2
generative world model
1
interactive world model
1
model-based reinforcement learning world model
0

Deep Analysis

Innovations

  • Introduction of the Eval-LLM-Drive dataset for evaluating MLLMs in driving scenarios
  • Development of the DriveSim simulator to support controlled evaluation
  • Systematic probing framework assessing MLLMs as world models from in-car camera perspectives

Methodology

The study conducts an experimental evaluation of various Multimodal Large Language Models (MLLMs), including GPT-4o, as world models for autonomous driving. It uses in-car camera perspectives and introduces the Eval-LLM-Drive dataset and DriveSim simulator to assess performance across four key areas: ego vehicle dynamics, interactions with road actors, trajectory planning, and open-set scene reasoning.

Key Results

MLLMs excel at interpreting individual images but struggle to synthesize coherent narratives across frames, leading to considerable inaccuracies in understanding ego vehicle dynamics, interactions with other road actors, trajectory planning, and open-set scene reasoning.

Limitations

  • The evaluation is limited to in-car camera perspectives and may not generalize to other sensor modalities or driving setups
  • The study identifies gaps in current MLLM capabilities but does not propose concrete improvements or solutions
  • The scope of the Eval-LLM-Drive dataset and DriveSim simulator (e.g., size, diversity) is not detailed in the abstract

Tags