Probing Multimodal LLMs as World Models for Driving
TLDR
MLLMs like GPT-4o struggle to understand dynamic driving scenes across frames, despite excelling at single images, highlighting gaps in world model capabilities.
Reasoning
Strengths: challenges common assumptions with systematic evaluation using a new dataset and simulator. Weaknesses: only identifies limitations without proposing solutions, and evaluation is limited to specific MLLMs.
Read-first score
Read-first score 60.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 24.
Field roles
Rank sensitivity
Stability: volatile; rank range: 645.
Keyword Scores
Deep Analysis
Innovations
- Introduction of the Eval-LLM-Drive dataset for evaluating MLLMs in driving scenarios
- Development of the DriveSim simulator to support controlled evaluation
- Systematic probing framework assessing MLLMs as world models from in-car camera perspectives
Methodology
The study conducts an experimental evaluation of various Multimodal Large Language Models (MLLMs), including GPT-4o, as world models for autonomous driving. It uses in-car camera perspectives and introduces the Eval-LLM-Drive dataset and DriveSim simulator to assess performance across four key areas: ego vehicle dynamics, interactions with road actors, trajectory planning, and open-set scene reasoning.
Key Results
MLLMs excel at interpreting individual images but struggle to synthesize coherent narratives across frames, leading to considerable inaccuracies in understanding ego vehicle dynamics, interactions with other road actors, trajectory planning, and open-set scene reasoning.
Limitations
- The evaluation is limited to in-car camera perspectives and may not generalize to other sensor modalities or driving setups
- The study identifies gaps in current MLLM capabilities but does not propose concrete improvements or solutions
- The scope of the Eval-LLM-Drive dataset and DriveSim simulator (e.g., size, diversity) is not detailed in the abstract