Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization

arXiv 2026 62.1 method

TLDR

CLAW learns continuous latent action world models from action-free videos via adversarial regularization and diffusion, enabling imitation learning and planning.

Reasoning

The paper presents a novel end-to-end self-supervised framework that jointly learns latent action representations and world models without action labels, demonstrating strong performance in imitation and planning tasks. However, the abstract lacks explicit mention of real-world experiments or limitations, and the reliance on diffusion-based generation may introduce computational overhead.

Read-first score

Read-first score 62.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 57.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
93.1

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.93061786

Topical relevance 29%
81.4

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 18%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=experiment,result

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridge

Rank sensitivity

Stability: volatile; rank range: 495.

Keyword Scores

world model
10
generative world model
9
video world model
9
world dynamics prediction
9
world simulator
7
model-based reinforcement learning world model
7
interactive world model
6

Deep Analysis

Innovations

  • End-to-end self-supervised learning of continuous latent actions and world model from action-free videos
  • Adversarial latent regularization to enforce structured and semantically meaningful latent action representations
  • Diffusion-based video generation for modeling rich predictive environment dynamics

Methodology

CLAW jointly trains a Latent Action Model and a world model using adversarial latent regularization and diffusion-based video generation, learning directly from action-free videos to infer continuous latent actions that explain environment transitions, without any action labels or annotations.

Key Results

Extensive experiments across diverse tasks and embodiments demonstrate that CLAW produces semantically meaningful latent action representations, supports effective action transfer, and enables planning and imitation from observation, outperforming existing methods.

Tags

world modelslatent actionsadversarial regularizationimitation learningvideo generationself-supervised learningRO