Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent

arXiv 25.1 2025 34 method

TLDR

UP-VLA unifies understanding and future prediction in VLA models, achieving 33% improvement on Calvin ABC-D and better real-world manipulation.

Reasoning

The paper introduces a novel training paradigm combining multi-modal understanding and future prediction, showing strong empirical results on both a benchmark and real-world tasks. However, it does not explicitly frame its prediction objective as a world model or simulator, limiting direct relevance to the specified keywords.

Read-first score

Read-first score 34, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 3.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,experiment,result

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Topical relevance 42%
4.3

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 198.

Keyword Scores

world dynamics prediction
2
world model
1
world simulator
0
generative world model
0
interactive world model
0
video world model
0
model-based reinforcement learning world model
0

Deep Analysis

Innovations

  • Unified training with both multi-modal understanding and future prediction objectives
  • Enhancing both high-level semantic comprehension and low-level spatial understanding for embodied control

Methodology

UP-VLA is a unified Vision-Language-Action model trained with both multi-modal understanding and future prediction objectives. It leverages pre-trained Vision-Language Models and is evaluated on the Calvin ABC-D benchmark and real-world manipulation tasks.

Key Results

UP-VLA achieves a 33% improvement on the Calvin ABC-D benchmark compared to the previous state-of-the-art method, and demonstrates improved success rates in real-world manipulation tasks requiring precise spatial information.

Tags