Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

arXiv 2026 44.4 benchmark

TLDR

Introduces Apple-PI, a benchmark to evaluate video generation models as law-grounded world simulators using classical mechanics tasks.

Reasoning

Strengths include a novel three-stage protocol for diagnosing reasoning bottlenecks and a hybrid evaluation suite combining subjective and objective measures. Weaknesses are the limited scope to classical mechanics and potential lack of diversity in physics tasks.

Read-first score

Read-first score 44.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 42.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 18%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,dataset,evaluation,metric

Topical relevance 29%
60

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Citation impact 18%
0

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. cited_by_count=0

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 445.

Keyword Scores

world model
9
world simulator
9
video world model
9
generative world model
8
world dynamics prediction
6
interactive world model
1
model-based reinforcement learning world model
0

Deep Analysis

Innovations

  • First benchmark to evaluate video models' physical law-grounded reasoning process, not just output plausibility
  • Orchard dataset: 400 videos, ten canonical classical mechanics tasks, with separation of single-law and multi-law tasks for confounder-free diagnosis and generalization probing
  • Three-stage scientific reasoning protocol (Perception, Formulation, Deduction) that uses chain-of-frames prompting on infographic-annotated first frames, treating generated video as visible reasoning trace
  • Hybrid evaluation suite combining MLLM-based subjective scoring with physics-law-grounded objective measures for stage-resolved failure diagnosis

Methodology

The benchmark comprises a dataset of 400 videos of classical mechanics tasks, a three-stage reasoning protocol using chain-of-frames prompting on infographic-annotated first frames to generate video as a reasoning trace, and a hybrid evaluation suite that combines MLLM subjective scoring with physics-law-grounded objective metrics. 11 video models are evaluated through this protocol, enabling stage-resolved diagnosis of perception, formulation, and deduction failures.

Key Results

The best video model achieves a score of only 0.473, indicating current models are far from reliable law-grounded world simulators. Stage-resolved analyses reveal a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap.

Tags