ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control
TLDR
A unified benchmark for evaluating adversarial robustness of world models in continuous control, testing attacks on policy, value, and latent dynamics.
Reasoning
The paper introduces ARB4WM, a comprehensive framework for assessing adversarial threats across multiple levels of world-model agents, addressing a clear gap in existing evaluations. Its strengths include systematic attack categorization and empirical evaluation on diverse tasks, but it is limited to Dreamer-style agents and continuous control domains, with input-level defenses showing limited effectiveness.
Read-first score
Read-first score 67.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 43.
Field roles
Rank sensitivity
Stability: volatile; rank range: 442.
Keyword Scores
Deep Analysis
Innovations
- Unified evaluation framework for pre-deployment robustness and risk assessment of world-model agents under visual perturbations
- Five white-box loss objectives targeting policy, value, and latent-dynamics levels
- Temporal attack modes: full-frame, half-sequence, and sparse-frame exposure
- Systematic evaluation of Dreamer-style agents across 20 tasks from MetaWorld and DeepMind Control Suite
Methodology
ARB4WM defines five white-box loss objectives across policy, value, and latent-dynamics levels, and studies their effects when combined with single-step or multi-step perturbation strategies and temporal attack modes (full-frame, half-sequence, sparse-frame). The framework evaluates four Dreamer-style agents on 20 continuous control tasks from MetaWorld and the DeepMind Control Suite under different loss objectives, perturbation strategies, and temporal attack modes.
Key Results
Attacks targeting value estimation, latent representations, and RSSM dynamics can be as damaging as direct policy disruption; early or frequent perturbations are especially harmful, and input-level defenses provide limited recovery under adaptive attacks.