MoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature Modulation
TLDR
MoWM combines latent and pixel-space world models via mixture-of-experts for improved embodied action planning, achieving SOTA on CALVIN and real tasks.
Reasoning
The paper introduces a novel hybrid framework that fuses motion-aware latent features with pixel-level details, addressing limitations of each approach. Strengths include strong empirical results on both simulated and real-world benchmarks, while weaknesses are not explicitly discussed in the abstract (e.g., computational cost or failure cases).
Read-first score
Read-first score 58.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 37.
Field roles
Rank sensitivity
Stability: volatile; rank range: 261.
Keyword Scores
Deep Analysis
Innovations
- Mixture-of-world-model framework that fuses representations from hybrid world models for embodied action planning
- Latent-to-pixel feature modulation to combine motion-aware latent features with pixel-space features
- Emphasis on action-relevant visual details by leveraging both compact motion-aware and fine-grained pixel-level representations
Methodology
MoWM combines motion-aware latent world model features with pixel-space features via latent-to-pixel feature modulation, enabling the model to emphasize action-relevant visual details for action decoding. The framework integrates hybrid world models and is evaluated on the CALVIN benchmark and real-world manipulation tasks.
Key Results
The method achieves state-of-the-art task success rates on CALVIN and real-world manipulation tasks, demonstrating superior generalization compared to prior approaches.