Imperfect World Models are Exploitable
TLDR
Defines model exploitation in RL, proves inevitability on large policy sets, and establishes formal bridge to reward hacking.
Reasoning
Strengths: Provides a novel formal definition of model exploitation and connects it to reward hacking theory, with rigorous proofs. Weaknesses: No empirical validation or real-world experiments; purely theoretical analysis.
Read-first score
Read-first score 47.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 23.
Field roles
Rank sensitivity
Stability: volatile; rank range: 297.
Keyword Scores
Deep Analysis
Innovations
- Novel definition of model exploitation in reinforcement learning, where a world model is exploitable if it implies a strict policy preference opposite to the true environment's transition model.
- General theory of reward hacking and model exploitation that proves exploitation is essentially unavoidable on large policy sets, with reward hacking as a special case.
- Relaxed notion of exploitation and derivation of a safe horizon within which exploitation can be avoided.
Methodology
The paper develops a formal theoretical framework, defining model exploitation analogously to reward hacking but showing the inevitability proof does not transfer. It then constructs a general theory that proves exploitation is unavoidable on large policy sets and derives conditions for unhackability in finite sets, finding no counterpart for exploitation. A relaxed notion is introduced to derive a safe horizon.
Key Results
The theory proves that model exploitation is essentially unavoidable on large policy sets, and the conditions that guarantee unhackability in finite policy sets do not preclude exploitation. A safe horizon is derived within which exploitation can be avoided under a relaxed definition.
Limitations
- Exploitation is unavoidable on large policy sets, limiting the possibility of safe planning.
- Conditions that prevent reward hacking in finite policy sets have no counterpart for model exploitation.
- The safe horizon derived for the relaxed notion may be restrictive in practice.