Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Interpreting Physics in Video World Models

arXiv 26.2 2026 51.9 method

TLDR

Interpretability study of video world models reveals a Physics Emergence Zone where physical variables become accessible, with direction encoded via circular geometry.

Reasoning

Strengths: Novel interpretability approach identifying a sharp transition layer and geometric encoding of direction. Weaknesses: Limited to encoder-based video transformers; no discussion of generative or interactive capabilities.

Read-first score

Read-first score 51.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 32.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=ablation,benchmark

Topical relevance 42%
45.7

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 430.

Keyword Scores

video world model
10
world model
8
world dynamics prediction
6
generative world model
3
world simulator
2
model-based reinforcement learning world model
2
interactive world model
1

Deep Analysis

Innovations

  • First interpretability study to directly examine physical representations inside large-scale video encoders
  • Identification of a sharp intermediate-depth transition (Physics Emergence Zone) where physical variables become accessible
  • Discovery that scalar quantities (speed, acceleration) are available from early layers, while motion direction becomes accessible only at the Physics Emergence Zone
  • Finding that direction is encoded through a high-dimensional population structure with circular geometry, requiring coordinated multi-feature intervention to control

Methodology

The study uses layerwise probing, subspace geometry, patch-level decoding, and targeted attention ablations to characterize where and how physical information is organized within encoder-based video transformers. It examines representations across architectures and decomposes motion into explicit variables (speed, acceleration, direction) to probe their accessibility at different layers.

Key Results

Physical variables become accessible at an intermediate-depth transition (Physics Emergence Zone), with scalar quantities available from early layers and motion direction only emerging at this zone; direction is encoded via a distributed, high-dimensional population structure with circular geometry, not as a factorized representation.

Tags