Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control
TLDR
Worldscape-MoE unifies heterogeneous action controls (camera, robot, hand) into a scalable Mixture-of-Experts world model using Diffusion Transformers.
Reasoning
The paper addresses fragmentation in video generation world models by proposing a unified framework that leverages shared physical regularities across different action modalities. Strengths include a novel MoE architecture and evidence that heterogeneous supervision improves individual control; weaknesses include limited domain scope and potential scalability challenges not fully addressed in the abstract.
Read-first score
Read-first score 47.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 58.
Field roles
Rank sensitivity
Stability: volatile; rank range: 570.
Keyword Scores
Deep Analysis
Innovations
- Modality-aware control injection for heterogeneous action interfaces
- Shared and control-specific experts in a Mixture-of-Experts architecture
- Progressive MoE tuning strategy for continual extension to new action modalities
Methodology
Worldscape-MoE is a Mixture-of-Experts world model built on Diffusion Transformers that uses modality-aware control injection, shared and control-specific experts, and a progressive tuning strategy to handle heterogeneous action modalities. It is evaluated on locomotion, robotic manipulation, and egocentric hand control tasks.
Key Results
Heterogeneous supervision improves individual control capabilities; Worldscape-MoE achieves strong results on WorldArena, improves locomotion and hand-control metrics, shows robust out-of-distribution generalization, and demonstrates scaling behavior as more control data and experts are added.