Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control

arXiv 2026 67.3 benchmark

TLDR

A unified benchmark for evaluating adversarial robustness of world models in continuous control, testing attacks on policy, value, and latent dynamics.

Reasoning

The paper introduces ARB4WM, a comprehensive framework for assessing adversarial threats across multiple levels of world-model agents, addressing a clear gap in existing evaluations. Its strengths include systematic attack categorization and empirical evaluation on diverse tasks, but it is limited to Dreamer-style agents and continuous control domains, with input-level defenses showing limited effectiveness.

Read-first score

Read-first score 67.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 43.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
94.8

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.94820973

Reproducibility 18%
81

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=code,github

Methodology quality 18%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,result

Topical relevance 29%
61.4

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchorReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 442.

Keyword Scores

world model
10
model-based reinforcement learning world model
9
world dynamics prediction
8
generative world model
7
world simulator
5
interactive world model
3
video world model
1

Deep Analysis

Innovations

  • Unified evaluation framework for pre-deployment robustness and risk assessment of world-model agents under visual perturbations
  • Five white-box loss objectives targeting policy, value, and latent-dynamics levels
  • Temporal attack modes: full-frame, half-sequence, and sparse-frame exposure
  • Systematic evaluation of Dreamer-style agents across 20 tasks from MetaWorld and DeepMind Control Suite

Methodology

ARB4WM defines five white-box loss objectives across policy, value, and latent-dynamics levels, and studies their effects when combined with single-step or multi-step perturbation strategies and temporal attack modes (full-frame, half-sequence, sparse-frame). The framework evaluates four Dreamer-style agents on 20 continuous control tasks from MetaWorld and the DeepMind Control Suite under different loss objectives, perturbation strategies, and temporal attack modes.

Key Results

Attacks targeting value estimation, latent representations, and RSSM dynamics can be as damaging as direct policy disruption; early or frequent perturbations are especially harmful, and input-level defenses provide limited recovery under adaptive attacks.

Tags

adversarial robustnessworld modelscontinuous controlbenchmarkvisual perturbationsAI