Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
TLDR
A large-scale multi-view visuotactile dataset for evaluating deformable object world models, comparing 2D video and 3D particle approaches.
Reasoning
The paper's strength lies in its massive real-world dataset and systematic comparison of world model paradigms, but it is limited to deformable objects and only provides a preliminary robot planning demonstration rather than a full method.
Read-first score
Read-first score 46.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 43.
Field roles
Rank sensitivity
Stability: volatile; rank range: 576.
Keyword Scores
Deep Analysis
Innovations
- Large-scale multi-view visuotactile dataset (Deform360) with 198 objects, 1,980 sequences, 215 hours, 41 cameras, and bimanual tactile grippers
- Novel markerless visuotactile 3D tracking pipeline for extracting dense geometry and motion
- Systematic comparison of 2D video models and 3D particle models for deformable world modeling
- Benchmark and insights into trade-offs between structural priors and scalability
Methodology
Deform360 dataset collected with 41 surround-view cameras and bimanual tactile grippers across 1,980 interaction sequences of 198 daily objects. A markerless visuotactile 3D tracking pipeline extracts dense geometry and motion. State-of-the-art 2D video and 3D particle world models are evaluated on this data, and a robot planning task demonstrates real-world applicability.
Key Results
The evaluation reveals trade-offs between 2D video models and 3D particle models regarding structural priors and scalability. A preliminary robot planning demonstration shows the dataset's potential for real-world deformable object manipulation.