Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

TRAP: Tail-aware Ranking Attack for World-Model Planning

arXiv 2026 51.8 method

TLDR

TRAP backdoor attack exploits long-tailed ranking vulnerability in world-model planning to hijack decisions via trajectory ranking manipulation.

Reasoning

Strengths: Novel attack vector targeting trajectory ranking in world models, with tail-aware loss and dual gating; experiments on DreamerV3 and TD-MPC2 show effectiveness. Weaknesses: Limited to simulated tasks; no real-world deployment or defense analysis.

Read-first score

Read-first score 51.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 42.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
70.4

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.70406801

Topical relevance 29%
60

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 18%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=evaluation,experiment

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FrontierBridge

Rank sensitivity

Stability: volatile; rank range: 275.

Keyword Scores

world model
10
model-based reinforcement learning world model
9
world dynamics prediction
8
generative world model
7
world simulator
6
interactive world model
2
video world model
0

Deep Analysis

Innovations

  • Identification of a backdoor vulnerability in world models rooted in the long-tailed ranking structure of imagined trajectories
  • Tail-aware ranking loss that focuses optimization on decision-critical trajectories
  • Dual gating mechanisms to stabilize optimization and regulate when and where the attack penalty is applied

Methodology

TRAP is a backdoor attack framework for world models that targets the ranking of imagined trajectories. It combines a tail-aware ranking loss to focus on decision-critical trajectories with dual gating mechanisms that stabilize optimization and regulate when and where the attack penalty is applied. Under trigger conditions, TRAP alters the relative ranking of imagined trajectories to redirect planning outcomes while largely maintaining the normal ranking structure on clean inputs.

Key Results

Experiments on DreamerV3 and TD-MPC2 across diverse tasks show that TRAP consistently induces sustained behavioral deviations and significant performance degradation.

Tags

backdoor attackworld modelplanningsecurityadversarial attackranking attackLGAI