Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

arXiv 2026 63.2 method, benchmark, application

TLDR

Wh0 uses generative video world models to produce scalable egocentric human-hand manipulation data, improving dexterous VLA model zero-shot success on real-world tasks.

Reasoning

The paper presents a novel framework that leverages generative world models to address data scarcity for dexterous manipulation, with strong empirical results on 18 real-world tasks. However, the approach is limited to egocentric human-hand data and may not fully address sim-to-real gaps inherent in generative video outputs.

Read-first score

Read-first score 63.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 47.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
92.7

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.92724431

Methodology quality 18%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=ablation,dataset

Topical relevance 29%
67.1

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
50

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=artifact,code,dataset,github

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 411.

Keyword Scores

world model
10
generative world model
10
video world model
9
world simulator
7
world dynamics prediction
6
interactive world model
3
model-based reinforcement learning world model
2

Deep Analysis

Innovations

  • Using generative video world models as scalable and controllable sources of egocentric human-hand manipulation data
  • Converting generated videos into robot-trainable supervision through hand motion reconstruction and visual editing
  • Co-training with limited real robot data to adapt pretrained VLA models for dexterous manipulation deployment
  • Demonstrating significant improvement in zero-shot success on unseen tasks (from 8.3% to 38.9%) across 18 real-world tasks

Methodology

Wh0 employs a generative world model conditioned on language, objects, and scenes to produce WM-H, a 50k-episode dataset of egocentric human-object interaction videos. The generated videos are then converted into robot-trainable supervision via hand motion reconstruction and visual editing. Finally, pretrained VLA models are co-trained on WM-H and a limited amount of real robot data to adapt to dexterous manipulation tasks.

Key Results

Across 18 real-world dexterous manipulation tasks, Wh0 improves zero-shot success on unseen tasks from 8.3% (model post-trained only on robot data) to 38.9% (with Wh0). Ablation studies confirm that scalable generation and scene/embodiment alignment are key drivers of performance gains.

Limitations

  • Dependence on the quality and realism of the generative world model, which may produce artifacts or unrealistic interactions
  • Potential inaccuracies in hand motion reconstruction from synthetic videos, affecting downstream robot training
  • Requirement for a limited amount of real robot data for co-training, which may still be costly to obtain
  • Generalization may be limited to tasks and scenes similar to those in the generated dataset

Tags

generative world modelsegocentric videodexterous manipulationdata generationhuman-object interactionVLA modelsRO