Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures
TLDR
Hand2World generates egocentric interaction videos from a single image using free-space hand gestures, with autoregressive framework and camera geometry embeddings.
Reasoning
Strengths: addresses key challenges in egocentric generation (distribution shift, camera-hand ambiguity, long videos) with novel conditioning and camera embeddings. Weaknesses: limited to egocentric view, may require further validation on diverse scenes.
Read-first score
Read-first score 65.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.
Field roles
Rank sensitivity
Stability: volatile; rank range: 144.
Keyword Scores
Deep Analysis
Innovations
- Occlusion-invariant hand conditioning based on projected 3D hand meshes
- Explicit camera geometry via per-pixel Plücker-ray embeddings
- Fully automated monocular annotation pipeline
- Distillation of a bidirectional diffusion model into a causal generator for arbitrary-length synthesis
Methodology
Hand2World is a unified autoregressive framework that uses occlusion-invariant hand conditioning via projected 3D hand meshes, injects explicit camera geometry through per-pixel Plücker-ray embeddings, and employs a fully automated monocular annotation pipeline. It distills a bidirectional diffusion model into a causal generator to enable arbitrary-length video synthesis.
Key Results
Experiments on three egocentric interaction benchmarks show substantial improvements in perceptual quality and 3D consistency while supporting camera control and long-horizon interactive generation.