Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence
TLDR
LingBot-Video is a Mixture-of-Experts video pretraining paradigm for embodied intelligence, using robot-oriented data and multi-dimensional reward to bridge digital creativity and physical actuation.
Reasoning
Strengths include a novel MoE architecture for efficiency, a data profiling engine with robot-oriented footage, and a multi-dimensional reward system for physical realism. Weaknesses are the lack of quantitative results and limited comparison to baselines in the abstract.
Read-first score
Read-first score 36, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 38.
Field roles
Rank sensitivity
Stability: volatile; rank range: 186.
Keyword Scores
Deep Analysis
Innovations
- Mixture-of-Experts (MoE) architecture for video pretraining for embodied intelligence, scaled from scratch
- Data profiling engine that augments internet videos with extensive robot-oriented footage (manipulation, navigation, egocentric) to instill understanding of actions and world dynamics
- Multi-dimensional reward system for alignment with physical rationality and task completion, beyond standard aesthetics and prompt-following criteria
- LingBot-Video as the first large-scale open-source MoE video foundation model bridging digital creativity and physical actuation
Methodology
The paper proposes a DiT-based video pretraining paradigm using a Mixture-of-Experts architecture to balance modeling capacity and inference efficiency. It constructs a data profiling engine that combines internet videos with robot-oriented footage (manipulation, navigation, egocentric). A multi-dimensional reward system enforces alignment with physical rationality and task completion during training.
Key Results
Comprehensive evaluations demonstrate the model's performance and efficiency as a video foundation model for embodied intelligence.