Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Signed Compression Progress on a Sealed Audit is Goodhart-Resistant

arXiv 2026 49.3 method, theory

TLDR

Proves that signed compression progress on a sealed audit is Goodhart-resistant, with telescoping reward and finite-audit bounds, validated by experiments.

Reasoning

The paper provides a rigorous theoretical proof and empirical validation on ARC-TGI, showing resistance to reward hacking. Strengths include clear formalization and mechanization; weakness is limited scope to sealed audits and specific model classes.

Read-first score

Read-first score 49.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 17.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
95.4

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.95367269

Methodology quality 18%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,experiment,result

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Topical relevance 29%
24.3

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 402.

Keyword Scores

world model
8
model-based reinforcement learning world model
3
world dynamics prediction
2
world simulator
1
generative world model
1
interactive world model
1
video world model
1

Deep Analysis

Innovations

  • Formal proof that signed decrease of a fixed sealed-audit loss yields telescoping reward equal to endpoint audit improvement, making it Goodhart-resistant.
  • Finite-audit bound with false-positive budget: cumulative empirical reward ≤ true audit improvement + 2Δ_n(F,δ), where Δ_n is the uniform audit deviation of the model class.
  • Identification of failure modes: clipping, scoring on agent's own stream, high-capacity model on reusable panel, or neural class with vacuous Δ_n.
  • Lean 4 mechanization of the structural core (telescoping, finite-audit bound, finite Gibbs, entropy floor).
  • Experiment suite on ARC-TGI grid-transformation generators with adaptive holdout attacks.

Methodology

The paper provides a theoretical analysis defining intrinsic reward as the signed decrease of a fixed sealed-audit loss, proving a telescoping property that ties cumulative reward to endpoint audit improvement. For finite audit panels, a bound using uniform audit deviation Δ_n is derived. Experiments are conducted on ARC-TGI grid-transformation generators with adaptive holdout attacks to test signed progress against clip-farming, stream leakage, noisy-TV curiosity, and reusable audits.

Key Results

Finite-audit deviation scales as n^{-0.527}; signed progress resists clip-farming, stream leakage, and noisy-TV curiosity; naive reusable audits are exploitable by black-box scalar feedback, while standard release defenses keep the attack below the 2Δ_n threshold.

Limitations

  • The guarantee disappears if progress is clipped, scored on the agent's own stream, exposed to a high-capacity model on a reusable panel, or applied to a neural class that makes Δ_n vacuous.
  • The result relies on a sealed audit and uniform deviation bound; it may not hold for non-finite model classes or without proper audit mechanisms.

Tags

intrinsic motivationcompression progressGoodhart's lawreward designreinforcement learningLGAIML