Attacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World Models
TLDR
Identifies trusted imagination in VLA policies as an attack surface, showing easy corruption but hard targeted steering, with a detector.
Reasoning
The paper introduces a novel attack surface on world models in imagine-then-act policies, supported by empirical evaluation on three targets and a parameter-free detector. Strengths include clear asymmetry analysis and thorough evaluation; weaknesses include limited perturbation model and adaptive attacker constraints.
Read-first score
Read-first score 52.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 33.
Field roles
Rank sensitivity
Stability: volatile; rank range: 425.
Keyword Scores
Deep Analysis
Innovations
- Identifying the trusted imagination (latent trajectory) as the exposed attack surface in imagine-then-act VLA policies, rather than the reactive policy.
- Asymmetry: corrupting the imagination is easy (off-manifold displacement), but steering it precisely to a specified on-manifold target is hard.
- A parameter-free denoiser detector that exploits the off-manifold property of corrupted imaginations.
Methodology
The attacker uses projected gradient descent through the fully differentiable observation-to-imagination map with L-infinity-bounded observation perturbations. The threat model is capability-based. Evaluation is performed on three models: RynnVLA-002, LingBot-VA, and LaDi-WM, using untargeted and targeted attacks, and a denoiser detector.
Key Results
Untargeted corruption is roughly 60x stronger than random and is detected at AUC 1.0. Targeted control remains bounded. An adaptive attacker evades detection only by forgoing corruption. The reactive policy remains robust to corrupted imagination, but a native imagination-driven MPC exhibits adversary-specific task failure (at epsilon=0.01, success 0.70 versus 0.05; Fisher p < 10^-4).
Limitations
- Targeted control remains bounded, indicating difficulty in achieving precise steering.
- An adaptive attacker can evade detection by forgoing corruption, limiting the detector's effectiveness.
- The attack only affects downstream oracles that rely on imagination (e.g., MPC), not the reactive policy itself.