WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation
TLDR
WristWorld generates wrist-view videos from anchor views using a 4D world model, improving VLA manipulation performance.
Reasoning
The paper introduces a novel two-stage method combining geometric reconstruction and video generation to bridge the anchor-wrist view gap, with strong empirical results on real robotic datasets. However, the approach is specialized to wrist-view generation and relies on a specific geometric model (VGGT), limiting generality.
Read-first score
Read-first score 57.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 35.
Field roles
Rank sensitivity
Stability: volatile; rank range: 310.
Keyword Scores
Deep Analysis
Innovations
- First 4D world model that generates wrist-view videos solely from anchor views, addressing the gap between abundant anchor views and scarce wrist views.
- Spatial Projection Consistency (SPC) Loss to estimate geometrically consistent wrist-view poses and 4D point clouds.
- Two-stage pipeline: Reconstruction (extending VGGT with SPC loss) and Generation (temporally coherent video synthesis from reconstructed perspective).
Methodology
WristWorld operates in two stages: (i) Reconstruction, which extends VGGT and incorporates a Spatial Projection Consistency (SPC) Loss to estimate geometrically consistent wrist-view poses and 4D point clouds; (ii) Generation, which employs a video generation model to synthesize temporally coherent wrist-view videos from the reconstructed perspective. The model is evaluated on Droid, Calvin, and Franka Panda datasets.
Key Results
WristWorld achieves state-of-the-art video generation with superior spatial consistency, and improves VLA performance by raising the average task completion length on Calvin by 3.81% and closing 42.4% of the anchor-wrist view gap.