HarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action Models
TLDR
HarmoWAM unifies predictive and reactive control via a world model with adaptive gating for generalizable and precise robot manipulation.
Reasoning
The paper clearly identifies a trade-off between two WAM paradigms and proposes a novel architecture with predictive and reactive experts and a gating mechanism. However, the abstract lacks explicit mention of real-world experiments or benchmarks, and the evaluation details are not provided, limiting assessment of empirical validity.
Read-first score
Read-first score 52.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 42.
Field roles
Rank sensitivity
Stability: volatile; rank range: 279.
Keyword Scores
Deep Analysis
Innovations
- Observation of fundamental trade-off between Imagine-then-Execute and Joint Modeling paradigms in World Action Models
- Proposal of HarmoWAM, an end-to-end WAM that unifies predictive and reactive control via a world model
- Design of two complementary action experts: predictive expert using latent dynamics for iterative action generation, and reactive expert inferring actions from predicted visual evolution
- Process-Adaptive Gating Mechanism for adaptive coordination between experts
- Demonstration of strong zero-shot generalization across six real-world robotic tasks with variations in background, position, and object semantics, outperforming prior VLA and WAM models by 33% and 29%
Methodology
HarmoWAM uses a world model to provide spatio-temporal physical priors that condition two complementary action experts: a predictive expert leveraging latent dynamics for iterative action generation, and a reactive expert directly inferring actions from predicted visual evolution. A Process-Adaptive Gating Mechanism automatically determines the timing and location of switching between these experts. The model is evaluated on six real-world robotic tasks across three training-unseen test environments covering variations in background, position, and object semantics.
Key Results
HarmoWAM achieves strong zero-shot generalization across unseen test environments, significantly outperforming prior state-of-the-art VLA models and WAMs by margins of 33% and 29%, respectively.