Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Is Your Driving World Model an All-Around Player?

arXiv 2026 63.6 benchmark, application

TLDR

Introduces WorldLens, a unified benchmark evaluating driving world models across pixel quality, geometry, closed-loop planning, and human perception.

Reasoning

The paper addresses a critical gap in evaluating driving world models beyond visual quality, proposing a comprehensive benchmark with human annotations and an auto-evaluator. Strengths include multi-dimensional assessment and real-world data; weaknesses are domain specificity and lack of detailed model performance breakdown.

Read-first score

Read-first score 63.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 52.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 18%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,dataset,evaluation,metric

Citation impact 18%
75.2

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.75159998

Topical relevance 29%
74.3

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 364.

Keyword Scores

world model
10
video world model
9
generative world model
8
world dynamics prediction
8
interactive world model
7
world simulator
6
model-based reinforcement learning world model
4

Deep Analysis

Innovations

  • WorldLens: a unified benchmark measuring world-model fidelity across five complementary aspects and 24 standardized dimensions, covering pixel quality, 4D geometry, closed-loop driving, and human perceptual alignment.
  • WorldLens-26K: a 26,808-entry human-annotated preference dataset pairing numerical scores with textual rationales.
  • WorldLens-Agent: a vision-language evaluator distilled from human judgments for scalable, explainable auto-assessment.

Methodology

The paper introduces WorldLens, a benchmark that evaluates driving world models across five aspects (pixel quality, 4D geometry, closed-loop driving, human perceptual alignment, and additional dimensions) totaling 24 standardized metrics. Six representative models are assessed using this benchmark, and a human-annotated preference dataset (WorldLens-26K) is collected to bridge algorithmic metrics with human perception. A vision-language agent (WorldLens-Agent) is then distilled from these annotations to enable automated evaluation.

Key Results

No existing model dominates across all axes: texture-rich models violate geometry, geometry-aware models lack behavioral fidelity, and even the strongest performers achieve only 2–3 out of 10 on human realism ratings.

Limitations

  • The benchmark is limited to six representative models, which may not cover the full diversity of driving world models.
  • Human annotations in WorldLens-26K may contain subjectivity and bias, potentially affecting the distilled agent's alignment.
  • The auto-evaluator (WorldLens-Agent) may not fully capture all nuances of human perception, especially in edge cases.

Tags

driving world modelsbenchmarkevaluationautonomous drivingvideo generationrealismCVRO