Sword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training
TLDR
Sword improves world model robustness for VLA policy post-training via style augmentation and dynamic latent bootstrapping, outperforming WoVR on LIBERO.
Reasoning
The paper directly addresses key weaknesses of world models as simulators—poor generalization and long-horizon error accumulation—with two novel techniques. However, it is evaluated only on the simulated LIBERO benchmark, limiting evidence of real-world applicability. The abstract clearly states the core contribution and experimental results.
Read-first score
Read-first score 63, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 50.
Field roles
Rank sensitivity
Stability: volatile; rank range: 372.
Keyword Scores
Deep Analysis
Innovations
- Structure-Guided Style Augmentation to disentangle visual textures from task-relevant dynamics
- Dynamic Latent Bootstrapping to maintain consistency between training and inference with low memory consumption
Methodology
Sword introduces a robust world model framework for VLA policy post-training. It employs Structure-Guided Style Augmentation to improve generalization by separating visual textures from dynamics, and Dynamic Latent Bootstrapping to align training and inference distributions while keeping memory usage low. The method is evaluated on the LIBERO benchmark against the WoVR baseline using metrics for generalization, generation quality, robustness, fidelity, and RL post-training success rate.
Key Results
Sword significantly outperforms the WoVR baseline on the LIBERO benchmark across all evaluated metrics, including generalization, generation quality, robustness, fidelity, and success rate of reinforcement-learning post-training for VLA models.
Limitations
- Evaluation is limited to the LIBERO benchmark, so generalization to other environments is not demonstrated.
- Comparison is only against a single baseline (WoVR), leaving relative performance against other world model approaches unknown.