Learning and Leveraging World Models in Visual Representation Learning
TLDR
Extends JEPA to predict photometric transformations in latent space, learning controllable representations that match or surpass prior self-supervised methods.
Reasoning
The paper introduces Image World Models (IWM), a novel extension of JEPA that predicts global photometric transformations, with clear contributions on conditioning, difficulty, and capacity. Strengths include a principled recipe and demonstrated adaptability via fine-tuning, but the abstract lacks explicit real-world benchmark details, though empirical comparisons are implied.
Read-first score
Read-first score 35.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 15.
Field roles
Rank sensitivity
Stability: volatile; rank range: 119.
Keyword Scores
Deep Analysis
Innovations
- Generalizing JEPA prediction task to a broader set of corruptions beyond missing parts
- Introducing Image World Models (IWM) that predict the effect of global photometric transformations in latent space
- Identifying three key aspects for learning performant IWMs: conditioning, prediction difficulty, and capacity
- Demonstrating that fine-tuned IWM world model matches or surpasses previous self-supervised methods
- Showing ability to control abstraction level of learned representations (invariant vs equivariant)
Methodology
The paper proposes Image World Models (IWM), an extension of Joint-Embedding Predictive Architecture (JEPA) that learns to predict the effect of global photometric transformations in latent space. The approach relies on three key aspects: conditioning, prediction difficulty, and capacity. The model is trained in a self-supervised manner and evaluated by fine-tuning on diverse tasks, comparing against prior self-supervised methods.
Key Results
A fine-tuned IWM world model matches or surpasses the performance of previous self-supervised methods. Additionally, learning with IWM allows control over the abstraction level of representations, enabling either invariant or equivariant features.