From Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World Models
TLDR
Argues world models must be physically grounded actionable simulators, not just video generators, using medical decision-making as a test case.
Reasoning
Strengths: Clearly identifies limitations of current world models (visual conflation) and proposes a reframing towards causal structure and constraints. Weaknesses: As a survey, it may lack novel empirical contributions; the medical stress test is mentioned but not detailed in abstract.
Read-first score
Read-first score 66, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 60.
Field roles
Rank sensitivity
Stability: volatile; rank range: 358.
Keyword Scores
Deep Analysis
Innovations
- Reframing world models as actionable simulators rather than visual engines, emphasizing physical grounding and causal structure.
- Proposing structured 4D interfaces, constraint-aware dynamics, and closed-loop evaluation as key components.
- Using medical decision-making as an epistemic stress test to demonstrate the necessity of counterfactual reasoning and intervention planning.
Methodology
This survey analyzes current world models, showing they fail under intervention and violate invariant constraints despite high-fidelity video generation. It proposes a reframing toward actionable simulators with structured 4D interfaces, constraint-aware dynamics, and closed-loop evaluation, using medical decision-making as a case study to illustrate the requirements.
Key Results
Modern world models frequently violate invariant constraints, fail under intervention, and break down in safety-critical decision-making, demonstrating that visual realism is an unreliable proxy for world understanding.