World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks
TLDR
Introduces World-Ego Modeling to decompose future evolution into world and ego components, with a new benchmark and model achieving SOTA.
Reasoning
Strengths include a novel conceptual paradigm, a new benchmark (HTEWorld) for hybrid tasks, and strong empirical results. Weaknesses are the lack of real-world validation and potential overfitting to the specific benchmark.
Read-first score
Read-first score 61.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 52.
Field roles
Rank sensitivity
Stability: volatile; rank range: 421.
Keyword Scores
Deep Analysis
Innovations
- World-Ego Modeling paradigm that decomposes future evolution into world and ego components
- Definition of world-ego boundary from motion-, semantic-, and intention-based perspectives
- Analysis of three disentanglement strategies: post-, pre-, and full disentanglement
- World-Ego Model (WEM) with an implicit separate world-ego planner and cascade-parallel mixture-of-experts (CP-MoE) diffusion generator
- HTEWorld benchmark: first long-horizon world modeling benchmark for hybrid navigation-manipulation tasks with 125K video clips and 300 multi-turn evaluation trajectories
Methodology
The paper introduces World-Ego Modeling, a paradigm that separates future evolution into world (instruction-agnostic scene regularities) and ego (robot-centric instruction-conditioned dynamics) components. It instantiates this as the World-Ego Model (WEM), which couples an implicit separate world-ego planner with a cascade-parallel mixture-of-experts (CP-MoE) diffusion generator. To evaluate, the authors construct HTEWorld, a benchmark containing 125K video clips (over 4.5M frames) with fine-grained action annotations and 300 multi-turn evaluation trajectories (over 2K instructions) for hybrid navigation-manipulation tasks.
Key Results
WEM achieves state-of-the-art performance on the HTEWorld benchmark while remaining competitive on existing manipulation-only benchmarks.