Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
TLDR
Introduces Apple-PI, a benchmark to evaluate video generation models as law-grounded world simulators using classical mechanics tasks.
Reasoning
Strengths include a novel three-stage protocol for diagnosing reasoning bottlenecks and a hybrid evaluation suite combining subjective and objective measures. Weaknesses are the limited scope to classical mechanics and potential lack of diversity in physics tasks.
Read-first score
Read-first score 44.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 42.
Field roles
Rank sensitivity
Stability: volatile; rank range: 445.
Keyword Scores
Deep Analysis
Innovations
- First benchmark to evaluate video models' physical law-grounded reasoning process, not just output plausibility
- Orchard dataset: 400 videos, ten canonical classical mechanics tasks, with separation of single-law and multi-law tasks for confounder-free diagnosis and generalization probing
- Three-stage scientific reasoning protocol (Perception, Formulation, Deduction) that uses chain-of-frames prompting on infographic-annotated first frames, treating generated video as visible reasoning trace
- Hybrid evaluation suite combining MLLM-based subjective scoring with physics-law-grounded objective measures for stage-resolved failure diagnosis
Methodology
The benchmark comprises a dataset of 400 videos of classical mechanics tasks, a three-stage reasoning protocol using chain-of-frames prompting on infographic-annotated first frames to generate video as a reasoning trace, and a hybrid evaluation suite that combines MLLM subjective scoring with physics-law-grounded objective metrics. 11 video models are evaluated through this protocol, enabling stage-resolved diagnosis of perception, formulation, and deduction failures.
Key Results
The best video model achieves a score of only 0.473, indicating current models are far from reliable law-grounded world simulators. Stage-resolved analyses reveal a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap.