Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

World Pilot: Steering Vision-Language-Action Models with World-Action Priors

arXiv 2026 53 method, system

TLDR

World Pilot augments VLA models with world-action priors via latent and action steering, achieving SOTA on zero-shot OOD and real-robot manipulation tasks.

Reasoning

The paper presents a novel dual-pathway steering mechanism that effectively integrates world model priors into VLA policies, with strong empirical results on both benchmarks and real robots. However, the abstract lacks discussion of limitations, computational overhead, or failure cases, and the reliance on pretrained world models may limit generalizability.

Read-first score

Read-first score 53, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 35.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
95.8

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.95832845

Topical relevance 29%
50

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 18%
50

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=github

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridge

Rank sensitivity

Stability: volatile; rank range: 441.

Keyword Scores

world model
8
video world model
7
world dynamics prediction
6
generative world model
5
model-based reinforcement learning world model
4
interactive world model
3
world simulator
2

Deep Analysis

Innovations

  • Augmenting Vision-Language-Action (VLA) models with world-action priors from a World-Action Model (WAM) via two complementary pathways: Latent Steering and Action Steering.
  • Latent Steering conditions the perception layer on a scene-evolution latent, providing an anticipated view of the scene.
  • Action Steering supplies an anticipated trajectory as a motion prior to the action generator.
  • The scene-evolution prior remains effective even when supplied by a video-pretrained world model that has not been action-post-trained.

Methodology

World Pilot augments a Vision-Language-Action (VLA) model with a World-Action Model (WAM) that provides two complementary priors: Latent Steering conditions the perception layer on a scene-evolution latent, and Action Steering supplies an anticipated trajectory as a motion prior. The framework is evaluated on the LIBERO-Plus zero-shot OOD benchmark and four real-robot manipulation tasks, achieving state-of-the-art success rates.

Key Results

World Pilot attains a state-of-the-art Total success rate of 84.7% on the LIBERO-Plus zero-shot OOD benchmark and the highest success rate on every real-robot setting across four manipulation tasks, with the largest margins under shifts in viewpoint, geometry, deformable state, and pose.

Tags

vision-language-action modelsrobotic manipulationworld modelsaction priorssteeringpolicy augmentationRO