IPR-1: Interactive Physical Reasoner
TLDR
IPR-1 combines world-model rollouts with a VLM policy and PhysCode action space to improve physical reasoning across 1000+ games, outperforming GPT-5.
Reasoning
The paper introduces a novel integration of world models and VLMs for interactive physical reasoning, supported by a large benchmark and strong empirical results. However, the abstract lacks discussion of limitations and the evaluation is limited to simulated games rather than real-world environments.
Read-first score
Read-first score 51.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 44.
Field roles
Rank sensitivity
Stability: volatile; rank range: 376.
Keyword Scores
Deep Analysis
Innovations
- IPR (Interactive Physical Reasoner) using world-model rollouts to score and reinforce a VLM's policy
- PhysCode: a physics-centric action code aligning semantic intent with dynamics to provide a shared action space for prediction and reasoning
- G2U (Game-to-Unseen) benchmark of 1,000+ heterogeneous games with significant visual domain gaps
Methodology
IPR uses world-model rollouts to score and reinforce a VLM's policy, and introduces PhysCode as a shared action space for prediction and reasoning. The model is pretrained on 1,000+ games from the G2U benchmark.
Key Results
IPR performs robustly on levels from primitive intuition to goal-driven reasoning, surpasses GPT-5 overall, and shows improved performance with more training games and interaction steps, including zero-shot transfer to unseen games.