Generalization of World Models under Environmental Variability for Vision-based Quadrotor Navigation
TLDR
Study of DreamerV3 world model robustness to environmental variability in vision-based quadrotor navigation, with real-world deployment and open-loop imagination.
Reasoning
The paper provides a systematic evaluation of world model generalization under environmental variability, using both simulation and real-world quadrotor experiments, which is a strength. However, it is limited to a single model architecture (DreamerV3) and a specific task (quadrotor navigation), potentially reducing generalizability.
Read-first score
Read-first score 59.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 50.
Field roles
Rank sensitivity
Stability: volatile; rank range: 399.
Keyword Scores
Deep Analysis
Innovations
- Systematic study of world model robustness to environmental variability using vision-based quadrotor navigation as a testbed
- Cross-environment validation spanning both Self-Supervised Learning (SSL) pretraining and Reinforcement Learning (RL) fine-tuning
- Real-world deployment with an open-loop scenario where the model navigates entirely in imagination after only 2.5 seconds of sensory input
- Identification of discrete latent size and training-sequence length as dominant factors governing world model quality
Methodology
The study uses DreamerV3-based world models trained under varying levels of environmental randomness. Models are evaluated via cross-environment validation across SSL pretraining and RL fine-tuning, and then deployed on a real quadrotor in unseen environments, including an open-loop test where the model receives 2.5s of real sensory input before navigating entirely in imagination over a 12m traverse.
Key Results
World model robustness during SSL pretraining is a strong predictor of sim-to-real transfer: every model that generalized well in cross-environment SSL validation deployed successfully in the real world (passing gaps as narrow as 0.67m), while the model that dominated simulation policy evaluation failed on the real platform. Discrete latent size and training-sequence length are identified as the dominant factors governing world model quality.