Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Imperfect World Models are Exploitable

arXiv 2026 47.4 theory

TLDR

Defines model exploitation in RL, proves inevitability on large policy sets, and establishes formal bridge to reward hacking.

Reasoning

Strengths: Provides a novel formal definition of model exploitation and connects it to reward hacking theory, with rigorous proofs. Weaknesses: No empirical validation or real-world experiments; purely theoretical analysis.

Read-first score

Read-first score 47.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 23.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
80.5

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.80450476

Methodology quality 18%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,result

Topical relevance 29%
32.9

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 297.

Keyword Scores

world model
9
model-based reinforcement learning world model
7
world dynamics prediction
3
world simulator
1
generative world model
1
interactive world model
1
video world model
1

Deep Analysis

Innovations

  • Novel definition of model exploitation in reinforcement learning, where a world model is exploitable if it implies a strict policy preference opposite to the true environment's transition model.
  • General theory of reward hacking and model exploitation that proves exploitation is essentially unavoidable on large policy sets, with reward hacking as a special case.
  • Relaxed notion of exploitation and derivation of a safe horizon within which exploitation can be avoided.

Methodology

The paper develops a formal theoretical framework, defining model exploitation analogously to reward hacking but showing the inevitability proof does not transfer. It then constructs a general theory that proves exploitation is unavoidable on large policy sets and derives conditions for unhackability in finite sets, finding no counterpart for exploitation. A relaxed notion is introduced to derive a safe horizon.

Key Results

The theory proves that model exploitation is essentially unavoidable on large policy sets, and the conditions that guarantee unhackability in finite policy sets do not preclude exploitation. A safe horizon is derived within which exploitation can be avoided under a relaxed definition.

Limitations

  • Exploitation is unavoidable on large policy sets, limiting the possibility of safe planning.
  • Conditions that prevent reward hacking in finite policy sets have no counterpart for model exploitation.
  • The safe horizon derived for the relaxed notion may be restrictive in practice.

Tags

reinforcement learningworld modelsmodel exploitationreward hackingpolicy optimizationtheoretical analysisAILG