Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

EgoExo-WM: Unlocking Exo Video for Ego World Models

arXiv 2026 56.7 method

TLDR

Using exocentric video to train egocentric world models via body pose extraction and video transformation improves prediction and planning.

Reasoning

The paper presents a novel method to leverage abundant exocentric video for egocentric world model training, showing improvements in prediction and planning. Strengths include a clear problem-motivated approach and empirical results; weaknesses are the lack of explicit real-world benchmark details in the abstract.

Read-first score

Read-first score 56.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 55.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Topical relevance 29%
78.6

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Citation impact 18%
76.7

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.76747133

Methodology quality 18%
50

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=none

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridge

Rank sensitivity

Stability: volatile; rank range: 479.

Keyword Scores

world model
10
video world model
9
world dynamics prediction
9
generative world model
8
interactive world model
8
model-based reinforcement learning world model
7
world simulator
4

Deep Analysis

Innovations

  • Extracting structured body pose from exocentric video as a representation of action
  • Transforming exocentric video to egocentric video using a human kinematics prior
  • Unlocking integration of in-the-wild exocentric data for egocentric world model training
  • Whole-body action-conditioned egocentric world models trained with converted data

Methodology

The method extracts structured body pose from exocentric video as a representation of action, then transforms the exocentric video to egocentric video using a human kinematics prior. This converted data is used to train whole-body action-conditioned egocentric world models.

Key Results

Training whole-body action-conditioned egocentric world models with the converted data significantly improves both prediction quality and downstream planning performance, specifically in inferring the sequence of body poses needed to achieve a visual goal state.

Limitations

  • No limitations are explicitly stated in the abstract.

Tags

egocentric visionexocentric videoworld modelsaction predictionhuman pose estimationvideo transformationCV