Representing Positional Information in Generative World Models for Object Manipulation
TLDR
Object manipulation capabilities are essential skills that set apart embodied agents engaging with the world, especially in the realm of robotics.
Reasoning
Fallback reasoning generated from available title and abstract metadata: Object manipulation capabilities are essential skills that set apart embodied agents engaging with the world, especially in the realm of robotics. The ability to predict outcomes of interactions with objects is paramount in this setting. While model-based control...
Read-first score
Read-first score 36.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals.
Field roles
Rank sensitivity
Stability: volatile; rank range: 146.
Deep Analysis
Innovations
- Identification of poor representation of positional information as the cause of underperformance in world models for object manipulation
- Introduction of two declinations for generative world models: position-conditioned (PCP) and latent-conditioned (LCP) policy learning
- LCP's use of object-centric latent representations that explicitly capture positional information, enabling multimodal goal specification via spatial coordinates or visual goal
Methodology
The paper proposes two approaches for generative world models: position-conditioned (PCP) and latent-conditioned (LCP) policy learning. LCP employs object-centric latent representations to capture object positional information for goal specification. The methods are evaluated across several manipulation environments against current model-based control baselines.
Key Results
The proposed methods show favorable performance compared to current model-based control approaches across several manipulation environments.