Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Testl
TLDR
Introduces WoW-World-Eval, a benchmark with 22 metrics to evaluate video foundation models as embodied world models, showing poor performance in planning and physical consistency.
Reasoning
Strengths include a comprehensive evaluation protocol with high human correlation (0.93) and clear focus on critical abilities. Weaknesses are limited data scope (609 robot manipulation samples) and poor model performance, which may indicate benchmark difficulty rather than model inadequacy.
Read-first score
Read-first score 68, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 54.
Field roles
Rank sensitivity
Stability: volatile; rank range: 147.
Keyword Scores
Deep Analysis
Innovations
- Introduction of the Embodied Turing Test benchmark (WoW-World-Eval) for evaluating video foundation models as predictive world models in Embodied AI
- Comprehensive evaluation protocol with 22 metrics covering five core abilities: perception, planning, prediction, generalization, and execution
- High Pearson correlation (0.93) between overall score and human preference, establishing a reliable foundation for the Human Turing Test
- Inverse Dynamic Model (IDM) Turing Test to assess execution accuracy in the real world
Methodology
The benchmark is built upon 609 robot manipulation data and evaluates five core abilities (perception, planning, prediction, generalization, execution) using a comprehensive protocol with 22 metrics. It also employs an Inverse Dynamic Model (IDM) to test the execution accuracy of video foundation models in real-world scenarios.
Key Results
Models achieve only 17.27 on long-horizon planning and at best 68.02 on physical consistency, indicating limited spatiotemporal consistency and physical reasoning. In the IDM Turing Test, most models collapse to approximately 0% success, while WoW maintains a 40.74% success rate.
Limitations
- Noticeable gap between generated videos and the real world
- Limited spatiotemporal consistency and physical reasoning in current models
- Low success rates in real-world execution accuracy (most models near 0%)