CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization
TLDR
CLAW learns continuous latent action world models from action-free videos via adversarial regularization and diffusion, enabling imitation learning and planning.
Reasoning
The paper presents a novel end-to-end self-supervised framework that jointly learns latent action representations and world models without action labels, demonstrating strong performance in imitation and planning tasks. However, the abstract lacks explicit mention of real-world experiments or limitations, and the reliance on diffusion-based generation may introduce computational overhead.
Read-first score
Read-first score 62.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 57.
Field roles
Rank sensitivity
Stability: volatile; rank range: 495.
Keyword Scores
Deep Analysis
Innovations
- End-to-end self-supervised learning of continuous latent actions and world model from action-free videos
- Adversarial latent regularization to enforce structured and semantically meaningful latent action representations
- Diffusion-based video generation for modeling rich predictive environment dynamics
Methodology
CLAW jointly trains a Latent Action Model and a world model using adversarial latent regularization and diffusion-based video generation, learning directly from action-free videos to infer continuous latent actions that explain environment transitions, without any action labels or annotations.
Key Results
Extensive experiments across diverse tasks and embodiments demonstrate that CLAW produces semantically meaningful latent action representations, supports effective action transfer, and enables planning and imitation from observation, outperforming existing methods.