Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Unified Video Action Model

arXiv 2025 32.7 method

TLDR

UVA jointly optimizes video and action predictions via joint latent representation and decoupled decoding for efficient robotics policy learning and dynamics modeling.

Reasoning

The paper introduces a unified framework that combines video generation and action prediction, demonstrating strong empirical results across multiple robotics tasks. However, the abstract lacks explicit details on real-world deployment and does not directly engage with the world model literature, limiting its novelty in that context.

Read-first score

Read-first score 32.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 10.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
40

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=experiment,result

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=github

Topical relevance 42%
14.3

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 96.

Keyword Scores

world dynamics prediction
4
video world model
3
world model
2
model-based reinforcement learning world model
1
world simulator
0
generative world model
0
interactive world model
0

Tags