WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World
TLDR
WorldLens is a full-spectrum benchmark for evaluating driving world models across visual, geometric, physical, and behavioral fidelity.
Reasoning
The paper introduces a comprehensive evaluation framework with a dataset and agent, addressing a clear gap in unified assessment. Strengths include multi-dimensional coverage and human alignment; weaknesses are not evident from the abstract alone, but the benchmark's scalability and generalizability remain to be seen.
Read-first score
Read-first score 50.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 51.
Field roles
Rank sensitivity
Stability: volatile; rank range: 672.
Keyword Scores
Deep Analysis
Innovations
- WorldLens benchmark covering five evaluation aspects: Generation, Reconstruction, Action-Following, Downstream Task, and Human Preference
- WorldLens-26K dataset of human-annotated videos with numerical scores and textual rationales
- WorldLens-Agent evaluation model distilled from human annotations for scalable, explainable scoring
Methodology
WorldLens is a full-spectrum benchmark that evaluates generative world models across five aspects: Generation, Reconstruction, Action-Following, Downstream Task, and Human Preference. To align objective metrics with human judgment, the authors construct WorldLens-26K, a large-scale dataset of human-annotated videos with numerical scores and textual rationales, and develop WorldLens-Agent, an evaluation model distilled from these annotations to enable scalable, explainable scoring.
Key Results
No existing world model excels universally across all dimensions: models with strong textures often violate physics, while geometry-stable ones lack behavioral fidelity.
Limitations
- Benchmark is specific to driving world models, limiting generalizability to other domains
- Human annotations may introduce subjective biases, and the distilled evaluation model may not perfectly capture all aspects of human judgment