LaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot Manipulation
TLDR
LaST-HD aligns human and robot demonstrations in a shared latent space using an action-conditioned world model for robot manipulation learning.
Reasoning
The paper introduces a novel paradigm for human-to-robot action learning by leveraging a world model to align cross-embodiment dynamics, which is a strength. However, the abstract lacks explicit details on real-world evaluation results and does not directly address several of the specified keywords beyond 'world model'.
Read-first score
Read-first score 39.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 11.
Field roles
Rank sensitivity
Stability: volatile; rank range: 321.
Keyword Scores
Deep Analysis
Innovations
- Extending reasoning-before-acting VLA by aligning human-hand and robot demonstrations in a shared latent reasoning space
- Training an auxiliary action-conditioned world model on unpaired human-hand and robot trajectories to synthesize unified latent targets
- Developing Out-of-Lab (OOL) Glove, a low-cost motion-capture glove for human-hand data collection
- Progressive mixed-to-human training recipe comprising mixed human-robot co-training and human-hand online correction post-training
Methodology
LaST-HD is a human-to-robot action learning paradigm that extends reasoning-before-acting VLA by aligning human-hand and robot demonstrations in a shared latent reasoning space. It trains an auxiliary action-conditioned world model on unpaired human-hand and robot trajectories to synthesize unified latent targets, then aligns cross-embodiment representations in this forward-dynamics space. The method uses the OOL Glove for data collection and employs a progressive mixed-to-human training recipe with mixed co-training and online correction.
Key Results
LaST-HD improves generalization to novel objects, scenes, and positions using only human-hand demonstrations. With online correction, it achieves over 90% accuracy using only 20 minutes of OOL glove data.