EgoExo-WM: Unlocking Exo Video for Ego World Models
TLDR
Using exocentric video to train egocentric world models via body pose extraction and video transformation improves prediction and planning.
Reasoning
The paper presents a novel method to leverage abundant exocentric video for egocentric world model training, showing improvements in prediction and planning. Strengths include a clear problem-motivated approach and empirical results; weaknesses are the lack of explicit real-world benchmark details in the abstract.
Read-first score
Read-first score 56.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 55.
Field roles
Rank sensitivity
Stability: volatile; rank range: 479.
Keyword Scores
Deep Analysis
Innovations
- Extracting structured body pose from exocentric video as a representation of action
- Transforming exocentric video to egocentric video using a human kinematics prior
- Unlocking integration of in-the-wild exocentric data for egocentric world model training
- Whole-body action-conditioned egocentric world models trained with converted data
Methodology
The method extracts structured body pose from exocentric video as a representation of action, then transforms the exocentric video to egocentric video using a human kinematics prior. This converted data is used to train whole-body action-conditioned egocentric world models.
Key Results
Training whole-body action-conditioned egocentric world models with the converted data significantly improves both prediction quality and downstream planning performance, specifically in inferring the sequence of body poses needed to achieve a visual goal state.
Limitations
- No limitations are explicitly stated in the abstract.