Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data
TLDR
Wh0 uses generative video world models to produce scalable egocentric human-hand manipulation data, improving dexterous VLA model zero-shot success on real-world tasks.
Reasoning
The paper presents a novel framework that leverages generative world models to address data scarcity for dexterous manipulation, with strong empirical results on 18 real-world tasks. However, the approach is limited to egocentric human-hand data and may not fully address sim-to-real gaps inherent in generative video outputs.
Read-first score
Read-first score 63.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 47.
Field roles
Rank sensitivity
Stability: volatile; rank range: 411.
Keyword Scores
Deep Analysis
Innovations
- Using generative video world models as scalable and controllable sources of egocentric human-hand manipulation data
- Converting generated videos into robot-trainable supervision through hand motion reconstruction and visual editing
- Co-training with limited real robot data to adapt pretrained VLA models for dexterous manipulation deployment
- Demonstrating significant improvement in zero-shot success on unseen tasks (from 8.3% to 38.9%) across 18 real-world tasks
Methodology
Wh0 employs a generative world model conditioned on language, objects, and scenes to produce WM-H, a 50k-episode dataset of egocentric human-object interaction videos. The generated videos are then converted into robot-trainable supervision via hand motion reconstruction and visual editing. Finally, pretrained VLA models are co-trained on WM-H and a limited amount of real robot data to adapt to dexterous manipulation tasks.
Key Results
Across 18 real-world dexterous manipulation tasks, Wh0 improves zero-shot success on unseen tasks from 8.3% (model post-trained only on robot data) to 38.9% (with Wh0). Ablation studies confirm that scalable generation and scene/embodiment alignment are key drivers of performance gains.
Limitations
- Dependence on the quality and realism of the generative world model, which may produce artifacts or unrealistic interactions
- Potential inaccuracies in hand motion reconstruction from synthetic videos, affecting downstream robot training
- Requirement for a limited amount of real robot data for co-training, which may still be costly to obtain
- Generalization may be limited to tasks and scenes similar to those in the generated dataset