Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation

arXiv 2026 63.7 method, benchmark, application

TLDR

WorldFly integrates world models with VLA for UAV navigation, using flow matching to predict future video and actions, outperforming baselines in dense urban environments.

Reasoning

The paper presents a novel world-model-based VLA framework for UAV navigation, with a challenging benchmark and strong empirical results. However, it is limited to simulated environments and does not address real-world deployment challenges or broader generalization.

Read-first score

Read-first score 63.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
93.5

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.93472496

Methodology quality 18%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,benchmark,evaluation,result

Topical relevance 29%
68.6

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 416.

Keyword Scores

world model
10
video world model
9
world dynamics prediction
9
generative world model
8
interactive world model
5
model-based reinforcement learning world model
4
world simulator
3

Deep Analysis

Innovations

  • Proposes WorldFly, a novel world-model-based Vision-Language-Action (VLA) framework for UAV navigation that integrates spatial imagination via a dual-branch coupled flow matching mechanism.
  • Introduces a dual-branch coupled flow matching mechanism that jointly generates future video predictions and navigation actions, explicitly guiding the agent's policy.
  • Constructs a challenging Urban Canyon Traversal Benchmark specifically designed to evaluate spatial understanding under severe occlusions and drastic viewpoint transitions.

Methodology

WorldFly employs a dual-branch coupled flow matching mechanism within a VLA framework to jointly predict future video frames and navigation actions, leveraging world models to imagine future states under partial observability. The model is evaluated on the newly introduced Urban Canyon Traversal Benchmark against baseline methods, focusing on performance in unseen environments.

Key Results

WorldFly outperforms other baselines on the Urban Canyon Traversal Benchmark, particularly in unseen environments, demonstrating the effectiveness of integrating world models into embodied aerial agents for robust decision-making.

Limitations

  • The benchmark is limited to urban canyon scenarios with severe occlusions and sharp turns, so generalization to other environments (e.g., open fields, forests) is not evaluated.
  • The reliance on future video prediction may introduce computational overhead, potentially limiting real-time deployment on resource-constrained UAVs.
  • The dual-branch coupled flow matching mechanism may struggle with highly dynamic or unpredictable environments not covered in the benchmark.

Tags

UAV navigationVision-Language-ActionWorld ModelUrban CanyonNavigationAI