BadWAM: When World-Action Models Dream Right but Act Wrong
TLDR
A framework for adversarial attacks that break alignment between imagined future and executed action in world-action models.
Reasoning
Strengths: novel attack surface (World-Action Drift) and two distinct attack types. Weaknesses: evaluation only on model variants, no real-world validation; limited scope of WAMs.
Read-first score
Read-first score 35.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 37.
Field roles
Rank sensitivity
Stability: volatile; rank range: 176.
Keyword Scores
Deep Analysis
Innovations
- Introduces BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks, a new class of adversarial attacks specific to world-action models (WAMs).
- Characterizes the attack surface along two criteria: attack strength and stealthiness.
- Proposes an action-only adversarial attack that directly drives the model toward task-failing actions.
- Proposes an imagination-preserving adversarial attack that induces harmful action shifts while keeping the predicted future close to the clean imagination, exposing stealthy WAM failures.
Methodology
BadWAM uses small visual perturbations to attack WAMs, instantiating two attack strategies: an action-only attack that disrupts task success, and an imagination-preserving attack that constrains future prediction drift while altering actions. Attacks are evaluated on different WAM variants under closed-loop execution, measuring task success rates.
Key Results
The action-only attack reduces model performance from 96.5% to 43.1% success. The imagination-preserving attack maintains strong attack performance with moderate future-preserving regularization, exposing a vulnerability where the model imagines a plausible future but executes desynchronized actions.