DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation
TLDR
DyWA jointly predicts future states and adapts to dynamics variations for robust non-prehensile manipulation using single-view point clouds.
Reasoning
The paper presents a novel framework that addresses key limitations of existing methods by unifying geometry, state, physics, and action modeling, achieving strong simulation and real-world results. However, the abstract lacks detailed comparison with other world model approaches and does not discuss potential failure cases or computational costs.
Read-first score
Read-first score 60.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 49.
Field roles
Rank sensitivity
Stability: volatile; rank range: 282.
Keyword Scores
Deep Analysis
Innovations
- Dynamics-adaptive World Action Model (DyWA) that jointly predicts future states while adapting to dynamics variations based on historical trajectories
- Unifying modeling of geometry, state, physics, and robot actions for robust policy learning under partial observability
- Using only single-view point cloud observations, reducing reliance on multi-view cameras and precise pose tracking
Methodology
DyWA is a framework that enhances action learning by jointly predicting future states and adapting to dynamics variations from historical trajectories. It unifies geometry, state, physics, and robot action modeling, and is evaluated using single-view point cloud observations in both simulation and real-world experiments against baselines.
Key Results
In simulation, DyWA improves success rate by 31.5% using single-view point cloud observations. In real-world experiments, it achieves an average success rate of 68%, demonstrating generalization across object geometries, varying table friction, and robustness in challenging scenarios such as half-filled water bottles and slippery surfaces.