UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
TLDR
UP-VLA unifies understanding and future prediction in VLA models, achieving 33% improvement on Calvin ABC-D and better real-world manipulation.
Reasoning
The paper introduces a novel training paradigm combining multi-modal understanding and future prediction, showing strong empirical results on both a benchmark and real-world tasks. However, it does not explicitly frame its prediction objective as a world model or simulator, limiting direct relevance to the specified keywords.
Read-first score
Read-first score 34, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 3.
Field roles
Rank sensitivity
Stability: volatile; rank range: 198.
Keyword Scores
Deep Analysis
Innovations
- Unified training with both multi-modal understanding and future prediction objectives
- Enhancing both high-level semantic comprehension and low-level spatial understanding for embodied control
Methodology
UP-VLA is a unified Vision-Language-Action model trained with both multi-modal understanding and future prediction objectives. It leverages pre-trained Vision-Language Models and is evaluated on the Calvin ABC-D benchmark and real-world manipulation tasks.
Key Results
UP-VLA achieves a 33% improvement on the Calvin ABC-D benchmark compared to the previous state-of-the-art method, and demonstrates improved success rates in real-world manipulation tasks requiring precise spatial information.