Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Sword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training

arXiv 2026 63 method

TLDR

Sword improves world model robustness for VLA policy post-training via style augmentation and dynamic latent bootstrapping, outperforming WoVR on LIBERO.

Reasoning

The paper directly addresses key weaknesses of world models as simulators—poor generalization and long-horizon error accumulation—with two novel techniques. However, it is evaluated only on the simulated LIBERO benchmark, limiting evidence of real-world applicability. The abstract clearly states the core contribution and experimental results.

Read-first score

Read-first score 63, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 50.

Methodology quality 18%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,benchmark,evaluation,experiment,metric

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
74.4

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.7440437

Topical relevance 29%
71.4

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 372.

Keyword Scores

world model
10
world simulator
9
generative world model
8
model-based reinforcement learning world model
8
world dynamics prediction
7
interactive world model
5
video world model
3

Deep Analysis

Innovations

  • Structure-Guided Style Augmentation to disentangle visual textures from task-relevant dynamics
  • Dynamic Latent Bootstrapping to maintain consistency between training and inference with low memory consumption

Methodology

Sword introduces a robust world model framework for VLA policy post-training. It employs Structure-Guided Style Augmentation to improve generalization by separating visual textures from dynamics, and Dynamic Latent Bootstrapping to align training and inference distributions while keeping memory usage low. The method is evaluated on the LIBERO benchmark against the WoVR baseline using metrics for generalization, generation quality, robustness, fidelity, and RL post-training success rate.

Key Results

Sword significantly outperforms the WoVR baseline on the LIBERO benchmark across all evaluated metrics, including generalization, generation quality, robustness, fidelity, and success rate of reinforcement-learning post-training for VLA models.

Limitations

  • Evaluation is limited to the LIBERO benchmark, so generalization to other environments is not demonstrated.
  • Comparison is only against a single baseline (WoVR), leaving relative performance against other world model approaches unknown.

Tags

world modelsvision-language-actionpolicy post-trainingrobustnessgeneralizationlatent bootstrappingCVAI