Is Your Driving World Model an All-Around Player?
TLDR
Introduces WorldLens, a unified benchmark evaluating driving world models across pixel quality, geometry, closed-loop planning, and human perception.
Reasoning
The paper addresses a critical gap in evaluating driving world models beyond visual quality, proposing a comprehensive benchmark with human annotations and an auto-evaluator. Strengths include multi-dimensional assessment and real-world data; weaknesses are domain specificity and lack of detailed model performance breakdown.
Read-first score
Read-first score 63.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 52.
Field roles
Rank sensitivity
Stability: volatile; rank range: 364.
Keyword Scores
Deep Analysis
Innovations
- WorldLens: a unified benchmark measuring world-model fidelity across five complementary aspects and 24 standardized dimensions, covering pixel quality, 4D geometry, closed-loop driving, and human perceptual alignment.
- WorldLens-26K: a 26,808-entry human-annotated preference dataset pairing numerical scores with textual rationales.
- WorldLens-Agent: a vision-language evaluator distilled from human judgments for scalable, explainable auto-assessment.
Methodology
The paper introduces WorldLens, a benchmark that evaluates driving world models across five aspects (pixel quality, 4D geometry, closed-loop driving, human perceptual alignment, and additional dimensions) totaling 24 standardized metrics. Six representative models are assessed using this benchmark, and a human-annotated preference dataset (WorldLens-26K) is collected to bridge algorithmic metrics with human perception. A vision-language agent (WorldLens-Agent) is then distilled from these annotations to enable automated evaluation.
Key Results
No existing model dominates across all axes: texture-rich models violate geometry, geometry-aware models lack behavioral fidelity, and even the strongest performers achieve only 2–3 out of 10 on human realism ratings.
Limitations
- The benchmark is limited to six representative models, which may not cover the full diversity of driving world models.
- Human annotations in WorldLens-26K may contain subjectivity and bias, potentially affecting the distilled agent's alignment.
- The auto-evaluator (WorldLens-Agent) may not fully capture all nuances of human perception, especially in edge cases.