RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail
TLDR
Adapts a foundation video world model to retail scenes, comparing egocentric vs exocentric adaptation using a new synchronized multi-view dataset.
Reasoning
The paper introduces a novel dataset and systematically compares viewpoint adaptation strategies, showing exocentric-only adaptation often outperforms combined. Strengths include rigorous evaluation with multiple metrics and statistical tests; weakness is limited scope to retail domain and no interactive or RL context.
Read-first score
Read-first score 40.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 41.
Field roles
Rank sensitivity
Stability: volatile; rank range: 244.
Keyword Scores
Deep Analysis
Innovations
- Introduction of RetailSMV, a synchronized multi-view retail video dataset from store-staff perspective
- Systematic comparison of egocentric vs. exocentric adaptation of a foundation video world model to retail scenes
- Finding that exocentric-only adaptation outperforms combined adaptation and that adding egocentric data hurts
- Identification of the near-horizon prediction window as the regime where adaptation is most beneficial
Methodology
They adapt a pretrained Cosmos3-Nano video diffusion model using Low-Rank Adaptation (LoRA) with three matched configurations (egocentric-only, exocentric-only, combined) on the RetailSMV dataset of 32,105 captioned clips, then evaluate on a 200-clip held-out test set using seven complementary metrics and a strict paired statistical protocol.
Key Results
Exocentric-only adaptation matches or exceeds combined adaptation on six of seven metrics and is significantly better on LPIPS, PSNR, and DreamSim; adding exocentric data to egocentric training helps, while adding egocentric data to exocentric training hurts; the adaptation gap is largest at the shortest rollout time.