WoMAP: World Models For Embodied Open-Vocabulary Object Localization
TLDR
WoMAP uses a latent world model with Gaussian Splatting and open-vocabulary detectors for zero-shot active object localization in robotics.
Reasoning
The paper presents a novel pipeline combining world models with scalable data generation and reward distillation, achieving strong empirical results in simulation and real hardware. However, the abstract lacks discussion of failure cases or limitations, and some keyword relevance is indirect.
Read-first score
Read-first score 59.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 43.
Field roles
Rank sensitivity
Stability: volatile; rank range: 199.
Keyword Scores
Deep Analysis
Innovations
- Gaussian Splatting-based real-to-sim-to-real pipeline for scalable data generation without expert demonstrations
- Distillation of dense reward signals from open-vocabulary object detectors
- Latent world model for dynamics and rewards prediction to ground high-level action proposals at inference time
Methodology
WoMAP uses a Gaussian Splatting-based real-to-sim-to-real pipeline to generate training data without expert demonstrations, distills dense rewards from open-vocabulary object detectors, and leverages a latent world model for dynamics and reward prediction to ground action proposals. The approach is evaluated in simulation and on hardware, with baselines including VLM and diffusion policy methods.
Key Results
WoMAP achieves more than 9x and 2x higher success rates compared to VLM and diffusion policy baselines, respectively, in zero-shot object localization tasks, and demonstrates strong generalization and sim-to-real transfer on a TidyBot.