Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Targeting World Models to Compromise Robot Learning Pipelines

arXiv 2026 59.4 method

TLDR

World models introduce a stealthy data poisoning attack vector in robot learning, enabling compromised policies despite safe training data.

Reasoning

The paper identifies a novel security vulnerability in world models used for robot learning, demonstrating effective attacks against state-of-the-art models. Strengths include a clear attack methodology and empirical validation; weaknesses include limited scope to specific attack types and potential lack of generality across all world model variants.

Read-first score

Read-first score 59.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 37.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
96.9

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.9689902

Methodology quality 18%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=dataset,evaluation,result

Topical relevance 29%
52.9

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 430.

Keyword Scores

world model
10
world dynamics prediction
7
world simulator
6
model-based reinforcement learning world model
5
generative world model
4
interactive world model
3
video world model
2

Deep Analysis

Innovations

  • Novel attack methods that inject malicious prompts or compromising transition dynamics into visibly safe teleoperated datasets, which are only activated when fed through a world model as input.
  • Demonstration of a full end-to-end backdoor on a downstream DRL policy and a proof-of-concept for the VLA setting using state-of-the-art action-conditioned and text-conditioned world models.

Methodology

The authors propose data poisoning attacks that target world models by embedding malicious prompts or transition dynamics into teleoperated datasets that appear safe. These poisoned datasets are then used as input to world models, which generate synthetic dangerous robot training trajectories. The attacks are evaluated against both action-conditioned and text-conditioned world models, with downstream effects measured on a DRL policy (full backdoor) and a VLA setting (proof-of-concept).

Key Results

The attacks successfully produce synthetic dangerous robot training trajectories, leading to unsafe or compromised robot policies. A full end-to-end backdoor is demonstrated on a downstream DRL policy, and a proof-of-concept is shown for the VLA setting.

Limitations

  • World models are vulnerable to stealthy data poisoning attacks that are difficult to detect using traditional methods.
  • The findings highlight the need for more secure world models and a reevaluation of their role in the robot learning supply chain, implying current defenses are insufficient.

Tags

world modelsrobot learningdata poisoningadversarial attackssupply chain securityROAICR