What-If World: A Causal Benchmark for General World Models in Embodied Scenarios
TLDR
Introduces a causal benchmark for evaluating video generation world models on physical intervention tasks, showing all models fail significantly.
Reasoning
Strengths: novel causal evaluation benchmark with real-world data and taxonomy; weaknesses: limited to two domains (driving and manipulation), and results show poor performance but no analysis of why.
Read-first score
Read-first score 64.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 58.
Field roles
Rank sensitivity
Stability: volatile; rank range: 443.
Keyword Scores
Deep Analysis
Innovations
- Causal benchmark for world models using paired prompts that vary one physical detail
- APEO rubric with four components (Adherence, Physics, Environment, Outcome)
- Taxonomy of six physical variables shared across driving and manipulation
Methodology
The benchmark consists of 319 prompt pairs built on real frames from nuScenes and DROID datasets, organized by a taxonomy of six physical variables. Each pair is scored using APEO, a four-part rubric checking video adherence to prompt, physical consistency, environment preservation, and correct outcome difference. Nine state-of-the-art video generation models are evaluated.
Key Results
No model exceeds 52% on the paired score, with open-source models clustering near 28%. Performance tracks visual prominence of the intervention rather than physics tractability, with visually subtle interventions scoring as low as 14.2% and visually pronounced ones reaching 40.4%.
Limitations
- Benchmark limited to driving and manipulation domains and six physical variables
- Only tests single-variable causal interventions, not multi-variable or more complex changes
- Relies on two specific datasets (nuScenes and DROID), which may not generalize to other embodied scenarios