World Pilot: Steering Vision-Language-Action Models with World-Action Priors
TLDR
World Pilot augments VLA models with world-action priors via latent and action steering, achieving SOTA on zero-shot OOD and real-robot manipulation tasks.
Reasoning
The paper presents a novel dual-pathway steering mechanism that effectively integrates world model priors into VLA policies, with strong empirical results on both benchmarks and real robots. However, the abstract lacks discussion of limitations, computational overhead, or failure cases, and the reliance on pretrained world models may limit generalizability.
Read-first score
Read-first score 53, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 35.
Field roles
Rank sensitivity
Stability: volatile; rank range: 441.
Keyword Scores
Deep Analysis
Innovations
- Augmenting Vision-Language-Action (VLA) models with world-action priors from a World-Action Model (WAM) via two complementary pathways: Latent Steering and Action Steering.
- Latent Steering conditions the perception layer on a scene-evolution latent, providing an anticipated view of the scene.
- Action Steering supplies an anticipated trajectory as a motion prior to the action generator.
- The scene-evolution prior remains effective even when supplied by a video-pretrained world model that has not been action-post-trained.
Methodology
World Pilot augments a Vision-Language-Action (VLA) model with a World-Action Model (WAM) that provides two complementary priors: Latent Steering conditions the perception layer on a scene-evolution latent, and Action Steering supplies an anticipated trajectory as a motion prior. The framework is evaluated on the LIBERO-Plus zero-shot OOD benchmark and four real-robot manipulation tasks, achieving state-of-the-art success rates.
Key Results
World Pilot attains a state-of-the-art Total success rate of 84.7% on the LIBERO-Plus zero-shot OOD benchmark and the highest success rate on every real-robot setting across four manipulation tasks, with the largest margins under shifts in viewpoint, geometry, deformable state, and pose.