Learning to Act without Actions
TLDR
Introduces LAPO, a method to recover latent actions from videos, enabling training of policies and world models without action labels, and pre-training on web videos.
Reasoning
The paper presents a novel method for extracting latent action information from videos, which is a significant step toward pre-training RL agents on unlabeled data. Strengths include the innovative approach and potential for scaling to web-scale video. Weaknesses are that experiments are only on procedurally-generated environments, and no real-world validation is mentioned.
Read-first score
Read-first score 39.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 29.
Field roles
Candidate
Rank sensitivity
Stability: volatile; rank range: 253.