When Does LeJEPA Learn a World Model?
TLDR
LeJEPA with Gaussian regularization provably recovers latent world variables, enabling optimal planning; validated on robotic control.
Reasoning
The paper provides strong theoretical guarantees for linear identifiability of latent variables under Gaussian priors, supported by experiments from 2D to high-dimensional robotic control. Weaknesses include limited scope to additive-noise transitions and lack of comparison to other world model methods.
Read-first score
Read-first score 55.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 34.
Field roles
Rank sensitivity
Stability: volatile; rank range: 337.
Keyword Scores
Deep Analysis
Innovations
- Proving that LeJEPA (alignment plus Gaussian regularization) achieves linear identifiability of latent variables under stationary additive-noise transitions.
- Establishing that the Gaussian distribution is the unique latent distribution for which this linear identifiability guarantee holds.
- Deriving an approximate identifiability result where the guarantee degrades gracefully.
- Demonstrating that linear, orthogonal identifiability enables optimal latent-space planning.
Methodology
The paper provides a theoretical proof using spectral decomposition, showing that alignment strictly penalizes nonlinearity, making the linear map optimal. It also proves a converse ruling out non-Gaussian alternatives. Experiments validate the theory across 2D to 1024-dimensional latents, including distributional ablations and pixel-based robotic control.
Key Results
LeJEPA linearly recovers the world's latent variables from nonlinear observations, with the Gaussian distribution being the unique latent distribution for which this guarantee holds. The approximate identifiability result degrades gracefully, and the method enables optimal latent-space planning.
Limitations
- The guarantee assumes latents evolve under stationary, additive-noise transitions.
- The linear identifiability result holds only for Gaussian latent distributions.
- The approximate identifiability guarantee degrades gracefully, implying non-perfect recovery in non-ideal conditions.
- The analysis is specific to LeJEPA and may not generalize to other world model architectures or non-stationary/non-additive dynamics.