Towards High-Consistency Embodied World Model with Multi-View Trajectory Videos
TLDR
MTV-World uses multi-view trajectory videos for high-consistency embodied world model visuomotor prediction.
Reasoning
The paper introduces a novel multi-view trajectory video control method to address spatial information loss in embodied world models, and proposes an auto-evaluation pipeline. Strengths include addressing a key limitation of low-level action translation and using multi-view for consistency. Weaknesses are the lack of explicit real-world validation details and potential complexity of multi-view setups.
Read-first score
Read-first score 63.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 52.
Field roles
Rank sensitivity
Stability: volatile; rank range: 255.
Keyword Scores
Deep Analysis
Innovations
- Using trajectory videos obtained through camera intrinsic/extrinsic parameters and Cartesian-space transformation as control signals instead of low-level actions
- Multi-view framework to compensate for spatial information loss from projecting 3D actions onto 2D images
- Auto-evaluation pipeline leveraging multimodal large models and referring video object segmentation models, with Jaccard Index for spatial consistency
Methodology
MTV-World forecasts future frames based on multi-view trajectory videos as input, conditioning on an initial frame per view. It employs camera intrinsic/extrinsic parameters and Cartesian-space transformation to convert low-level actions into trajectory videos, and uses a multi-view framework to mitigate spatial information loss from 2D projection.
Key Results
Extensive experiments demonstrate that MTV-World achieves precise control execution and accurate physical interaction modeling in complex dual-arm scenarios.