TRAP: Tail-aware Ranking Attack for World-Model Planning
TLDR
TRAP backdoor attack exploits long-tailed ranking vulnerability in world-model planning to hijack decisions via trajectory ranking manipulation.
Reasoning
Strengths: Novel attack vector targeting trajectory ranking in world models, with tail-aware loss and dual gating; experiments on DreamerV3 and TD-MPC2 show effectiveness. Weaknesses: Limited to simulated tasks; no real-world deployment or defense analysis.
Read-first score
Read-first score 51.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 42.
Field roles
Rank sensitivity
Stability: volatile; rank range: 275.
Keyword Scores
Deep Analysis
Innovations
- Identification of a backdoor vulnerability in world models rooted in the long-tailed ranking structure of imagined trajectories
- Tail-aware ranking loss that focuses optimization on decision-critical trajectories
- Dual gating mechanisms to stabilize optimization and regulate when and where the attack penalty is applied
Methodology
TRAP is a backdoor attack framework for world models that targets the ranking of imagined trajectories. It combines a tail-aware ranking loss to focus on decision-critical trajectories with dual gating mechanisms that stabilize optimization and regulate when and where the attack penalty is applied. Under trigger conditions, TRAP alters the relative ranking of imagined trajectories to redirect planning outcomes while largely maintaining the normal ranking structure on clean inputs.
Key Results
Experiments on DreamerV3 and TD-MPC2 across diverse tasks show that TRAP consistently induces sustained behavioral deviations and significant performance degradation.