Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios

arXiv 2026 64.8 benchmark

TLDR

Introduces a causal benchmark for evaluating video generation world models on physical intervention tasks, showing all models fail significantly.

Reasoning

Strengths: novel causal evaluation benchmark with real-world data and taxonomy; weaknesses: limited to two domains (driving and manipulation), and results show poor performance but no analysis of why.

Read-first score

Read-first score 64.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 58.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
87.6

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.87646789

Topical relevance 29%
82.9

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 18%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,dataset

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 443.

Keyword Scores

world model
10
world simulator
9
video world model
9
world dynamics prediction
9
generative world model
8
model-based reinforcement learning world model
7
interactive world model
6

Deep Analysis

Innovations

  • Causal benchmark for world models using paired prompts that vary one physical detail
  • APEO rubric with four components (Adherence, Physics, Environment, Outcome)
  • Taxonomy of six physical variables shared across driving and manipulation

Methodology

The benchmark consists of 319 prompt pairs built on real frames from nuScenes and DROID datasets, organized by a taxonomy of six physical variables. Each pair is scored using APEO, a four-part rubric checking video adherence to prompt, physical consistency, environment preservation, and correct outcome difference. Nine state-of-the-art video generation models are evaluated.

Key Results

No model exceeds 52% on the paired score, with open-source models clustering near 28%. Performance tracks visual prominence of the intervention rather than physics tractability, with visually subtle interventions scoring as low as 14.2% and visually pronounced ones reaching 40.4%.

Limitations

  • Benchmark limited to driving and manipulation domains and six physical variables
  • Only tests single-variable causal interventions, not multi-variable or more complex changes
  • Relies on two specific datasets (nuScenes and DROID), which may not generalize to other embodied scenarios

Tags

video generationworld modelscausal reasoningbenchmarkphysical plausibilityembodied scenariosCV