Olaf-World: Orienting Latent Actions for Video World Modeling
TLDR
Olaf-World learns latent actions from unlabeled video by aligning them to temporal feature differences, enabling zero-shot action transfer for video world models.
Reasoning
The paper introduces a novel alignment objective (SeqΔ-REPA) to learn structured latent action spaces from unlabeled video, addressing the scarcity of action labels. Strengths include strong zero-shot transfer and data-efficient adaptation, but the abstract lacks explicit discussion of limitations or real-world deployment details.
Read-first score
Read-first score 50, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 53.
Field roles
Rank sensitivity
Stability: volatile; rank range: 467.
Keyword Scores
Deep Analysis
Innovations
- Key insight that semantic effects of actions are observable and can serve as a shared reference for aligning latent actions across contexts
- SeqΔ-REPA: a sequence-level control-effect alignment objective that anchors integrated latent action to temporal feature differences from a frozen self-supervised video encoder
- Olaf-World pipeline for pretraining action-conditioned video world models from large-scale passive video
Methodology
The paper introduces SeqΔ-REPA, a sequence-level control-effect alignment objective that aligns integrated latent actions with temporal feature differences from a frozen self-supervised video encoder. This objective is used in the Olaf-World pipeline to pretrain action-conditioned video world models from large-scale passive video without action labels.
Key Results
The method learns a more structured latent action space, leading to stronger zero-shot action transfer and more data-efficient adaptation to new control interfaces compared to state-of-the-art baselines.