PhyGround: Benchmarking Physical Reasoning in Generative World Models
TLDR
PhyGround benchmarks physical reasoning in generative world models via 250 prompts, 13 physical laws, and a large-scale human study with automated evaluator.
Reasoning
The paper's strength lies in its rigorous benchmark design with per-law diagnostics and a large-scale, quality-controlled human study (459 annotators, high split-half reliability). Weaknesses include potential limited coverage of physical laws and lack of explicit discussion on generalizability beyond the selected laws.
Read-first score
Read-first score 59.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 46.
Field roles
Rank sensitivity
Stability: volatile; rank range: 337.
Keyword Scores
Deep Analysis
Innovations
- Criteria-grounded benchmark with 250 curated prompts, each augmented with an expected physical outcome
- Taxonomy of 13 physical laws across solid-body mechanics, fluid dynamics, and optics, operationalized through observable sub-questions for per-law diagnostics
- Large-scale, quality-controlled human study grounded on social science lab experiment design, with 459 annotators, 5,796 complete annotations, and over 37.4K fine-grained labels
- PhyJudge-9B, an open physics-specialized VLM judge with substantially lower aggregate relative bias (3.3%) compared to Gemini-3.1-Pro (16.6%)
Methodology
PhyGround consists of 250 curated prompts, each paired with an expected physical outcome, and a taxonomy of 13 physical laws (solid-body mechanics, fluid dynamics, optics) operationalized through observable sub-questions. Eight modern video generation models are evaluated via a large-scale human study with 459 annotators providing 5,796 complete annotations and over 37.4K fine-grained labels, followed by quality control that yields high split-half model-ranking correlations (Spearman's rho > 0.90). An open physics-specialized VLM judge, PhyJudge-9B, is released for reproducible automated evaluation.
Key Results
The human annotations after quality control exhibit high split-half model-ranking correlations (Spearman's rho > 0.90). PhyJudge-9B achieves substantially lower aggregate relative bias (3.3%) than Gemini-3.1-Pro (16.6%).