Interpreting Physics in Video World Models
TLDR
Interpretability study of video world models reveals a Physics Emergence Zone where physical variables become accessible, with direction encoded via circular geometry.
Reasoning
Strengths: Novel interpretability approach identifying a sharp transition layer and geometric encoding of direction. Weaknesses: Limited to encoder-based video transformers; no discussion of generative or interactive capabilities.
Read-first score
Read-first score 51.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 32.
Field roles
Rank sensitivity
Stability: volatile; rank range: 430.
Keyword Scores
Deep Analysis
Innovations
- First interpretability study to directly examine physical representations inside large-scale video encoders
- Identification of a sharp intermediate-depth transition (Physics Emergence Zone) where physical variables become accessible
- Discovery that scalar quantities (speed, acceleration) are available from early layers, while motion direction becomes accessible only at the Physics Emergence Zone
- Finding that direction is encoded through a high-dimensional population structure with circular geometry, requiring coordinated multi-feature intervention to control
Methodology
The study uses layerwise probing, subspace geometry, patch-level decoding, and targeted attention ablations to characterize where and how physical information is organized within encoder-based video transformers. It examines representations across architectures and decomposes motion into explicit variables (speed, acceleration, direction) to probe their accessibility at different layers.
Key Results
Physical variables become accessible at an intermediate-depth transition (Physics Emergence Zone), with scalar quantities available from early layers and motion direction only emerging at this zone; direction is encoded via a distributed, high-dimensional population structure with circular geometry, not as a factorized representation.