Predicting Closed-Loop Performance of Latent World Models: Offline Checkpoint Selection for MPC and Model-Based RL Under Non-Markovian Rewards in LunarLander
TLDR
Proposes offline checkpoint selection metrics for latent world models, using Reward Observability Fraction to predict closed-loop performance in LunarLander.
Reasoning
Strengths include novel diagnostic metrics (ROF, CROF) for checkpoint selection and empirical validation showing improved performance. Weaknesses are limited scope (LunarLander with shaped rewards) and potential lack of generalizability.
Read-first score
Read-first score 40.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 26.
Field roles
FrontierReproducibility anchor
Rank sensitivity
Stability: volatile; rank range: 460.