Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Testl

arXiv 26.1 2026 68 benchmark

TLDR

Introduces WoW-World-Eval, a benchmark with 22 metrics to evaluate video foundation models as embodied world models, showing poor performance in planning and physical consistency.

Reasoning

Strengths include a comprehensive evaluation protocol with high human correlation (0.93) and clear focus on critical abilities. Weaknesses are limited data scope (609 robot manipulation samples) and poor model performance, which may indicate benchmark difficulty rather than model inadequacy.

Read-first score

Read-first score 68, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 54.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,metric

Topical relevance 42%
77.1

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 147.

Keyword Scores

world model
10
video world model
10
generative world model
9
world dynamics prediction
8
interactive world model
7
world simulator
6
model-based reinforcement learning world model
4

Deep Analysis

Innovations

  • Introduction of the Embodied Turing Test benchmark (WoW-World-Eval) for evaluating video foundation models as predictive world models in Embodied AI
  • Comprehensive evaluation protocol with 22 metrics covering five core abilities: perception, planning, prediction, generalization, and execution
  • High Pearson correlation (0.93) between overall score and human preference, establishing a reliable foundation for the Human Turing Test
  • Inverse Dynamic Model (IDM) Turing Test to assess execution accuracy in the real world

Methodology

The benchmark is built upon 609 robot manipulation data and evaluates five core abilities (perception, planning, prediction, generalization, execution) using a comprehensive protocol with 22 metrics. It also employs an Inverse Dynamic Model (IDM) to test the execution accuracy of video foundation models in real-world scenarios.

Key Results

Models achieve only 17.27 on long-horizon planning and at best 68.02 on physical consistency, indicating limited spatiotemporal consistency and physical reasoning. In the IDM Turing Test, most models collapse to approximately 0% success, while WoW maintains a 40.74% success rate.

Limitations

  • Noticeable gap between generated videos and the real world
  • Limited spatiotemporal consistency and physical reasoning in current models
  • Low success rates in real-world execution accuracy (most models near 0%)

Tags