PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
TLDR
PhysicsMind is a unified benchmark with real and simulated environments to evaluate physical reasoning and prediction in VLMs and world models.
Reasoning
The paper introduces a novel benchmark that combines real and simulated data to test physical law consistency, addressing fragmentation in existing benchmarks. Its strengths include a focused evaluation on three canonical physics principles and both VQA and video generation tasks. Weaknesses are that it only evaluates existing models without proposing new methods, and the data is not yet publicly available.
Read-first score
Read-first score 59.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 36.
Field roles
Rank sensitivity
Stability: volatile; rank range: 383.
Keyword Scores
Deep Analysis
Innovations
- Introduces PhysicsMind, a unified benchmark combining real and simulation environments for evaluating physical reasoning and generation in foundational VLMs and world models.
- Focuses on three canonical physics principles: Center of Mass, Lever Equilibrium, and Newton's First Law.
- Defines two complementary tasks: VQA (reasoning about physical quantities from images/videos) and Video Generation (evaluating trajectory adherence to physical laws).
- Provides a systematic evaluation across a broad range of recent MLLMs and video generation models, revealing reliance on appearance heuristics and frequent violations of basic mechanics.
Methodology
PhysicsMind is a benchmark comprising both real and simulated environments, designed to test law-consistent reasoning and generation over three canonical principles: Center of Mass, Lever Equilibrium, and Newton's First Law. It includes two main tasks: VQA tasks that assess models' ability to reason about physical quantities from images or short videos, and Video Generation tasks that evaluate whether predicted motion trajectories obey the same physical constraints as ground truth. The benchmark evaluates a broad range of recent multimodal large language models and video generation models, using metrics that compare model outputs to ground-truth physical quantities and trajectories.
Key Results
Evaluated models are found to rely on appearance heuristics while often violating basic mechanics, indicating that current scaling and training are insufficient for robust physical understanding.
Limitations
- Benchmark is limited to three canonical physics principles (Center of Mass, Lever Equilibrium, Newton's First Law), which may not cover the full spectrum of physical reasoning.
- Data and environments are not yet publicly available (to be released upon acceptance), limiting immediate reproducibility and external validation.
- The evaluation focuses on specific tasks (VQA and video generation) and may not capture deeper or more complex physical interactions present in real-world scenarios.