Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

arXiv 2026 36 method

TLDR

LingBot-Video is a Mixture-of-Experts video pretraining paradigm for embodied intelligence, using robot-oriented data and multi-dimensional reward to bridge digital creativity and physical actuation.

Reasoning

Strengths include a novel MoE architecture for efficiency, a data profiling engine with robot-oriented footage, and a multi-dimensional reward system for physical realism. Weaknesses are the lack of quantitative results and limited comparison to baselines in the abstract.

Read-first score

Read-first score 36, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 38.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Topical relevance 29%
54.3

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 18%
50

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=evaluation

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Citation impact 18%
0

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. cited_by_count=0

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 186.

Keyword Scores

video world model
8
generative world model
7
world dynamics prediction
7
world model
6
interactive world model
5
world simulator
3
model-based reinforcement learning world model
2

Deep Analysis

Innovations

  • Mixture-of-Experts (MoE) architecture for video pretraining for embodied intelligence, scaled from scratch
  • Data profiling engine that augments internet videos with extensive robot-oriented footage (manipulation, navigation, egocentric) to instill understanding of actions and world dynamics
  • Multi-dimensional reward system for alignment with physical rationality and task completion, beyond standard aesthetics and prompt-following criteria
  • LingBot-Video as the first large-scale open-source MoE video foundation model bridging digital creativity and physical actuation

Methodology

The paper proposes a DiT-based video pretraining paradigm using a Mixture-of-Experts architecture to balance modeling capacity and inference efficiency. It constructs a data profiling engine that combines internet videos with robot-oriented footage (manipulation, navigation, egocentric). A multi-dimensional reward system enforces alignment with physical rationality and task completion during training.

Key Results

Comprehensive evaluations demonstrate the model's performance and efficiency as a video foundation model for embodied intelligence.

Tags