Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation

arXiv 26.2 2026 64.5 method, application

TLDR

Proposes a video world model framework for instruction-driven robot manipulation using video generation and adversarial distillation for fast, accurate future prediction.

Reasoning

The paper addresses a clear gap in robotic manipulation by integrating video generation with action models, showing strong empirical results. However, the abstract lacks specific experimental details and does not explicitly confirm real-world benchmarks, though real observations are mentioned.

Read-first score

Read-first score 64.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 49.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Topical relevance 42%
70

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 25%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,experiment,result

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 205.

Keyword Scores

world model
10
video world model
10
world dynamics prediction
9
generative world model
8
interactive world model
6
world simulator
4
model-based reinforcement learning world model
2

Deep Analysis

Innovations

  • Framework for fast and predictive video-conditioned action in instruction-driven robot manipulation
  • Selection and adaptation of a robust video generation model for reliable future predictions
  • Adversarial distillation for fast few-step video generation
  • Action model that leverages both generated videos and real observations to correct spatial errors

Methodology

The framework first selects and adapts a robust video generation model to ensure reliable future predictions. It then applies adversarial distillation to enable fast, few-step video generation. Finally, it trains an action model that uses both generated videos and real observations to correct spatial errors, enabling precise manipulation.

Key Results

The method produces temporally coherent and spatially accurate video predictions that directly support precise manipulation, achieving significant improvements in embodiment consistency, spatial referring ability, and task completion over existing baselines.

Tags