AI Can Learn Scientific Taste
TLDR
Proposes RLCF to train AI to judge and propose high-impact research ideas, showing AI can learn scientific taste.
Reasoning
Strengths include a novel training paradigm using community feedback (citations) for preference modeling and alignment, with strong empirical results outperforming SOTA LLMs. Weaknesses include reliance on citation-based proxy for impact, which may not fully capture scientific taste, and limited scope to idea generation rather than full research automation.
Read-first score
Read-first score 51.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 60.
Field roles
Rank sensitivity
Stability: volatile; rank range: 62.
Keyword Scores
Deep Analysis
Innovations
- Reinforcement Learning from Community Feedback (RLCF) paradigm for learning scientific taste from large-scale community signals
- Formulation of scientific taste learning as a preference modeling and alignment problem
- Scientific Judge model trained on 700K field- and time-matched high- vs. low-citation paper pairs to evaluate research ideas
- Scientific Thinker policy model trained via preference alignment using Scientific Judge as a reward model to generate high-impact research ideas
Methodology
The authors propose RLCF, first training a Scientific Judge on 700K matched pairs of high- and low-citation papers to model scientific preference, then using it as a reward model to align a policy model, Scientific Thinker, to propose ideas with high potential impact.
Key Results
Scientific Judge outperforms state-of-the-art LLMs (GPT-5.2, Gemini 3 Pro) and generalizes to future years, unseen fields, and peer-review preferences; Scientific Thinker generates research ideas with higher potential impact than baselines.