基于自监督预测的好奇心驱动探索
TLDR
好奇心作为自监督学习特征空间中的预测误差,实现了稀疏奖励环境下的探索。
评分理由
The paper introduces a novel curiosity formulation that scales to high-dimensional states and ignores irrelevant features, but its evaluation is limited to two game environments and lacks comparisons to other exploration methods.
Read-first 评分解释
综合优先阅读分 20.2,由主题、引用、图谱、方法、可复现性和近期性等信号加权得到。 原始总分保留为 9。
研究版图角色
候选论文
排序敏感性
稳定性:volatile;排名波动范围:17。