Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

AI Can Learn Scientific Taste

arXiv 2026 51.7 method

TLDR

Proposes RLCF to train AI to judge and propose high-impact research ideas, showing AI can learn scientific taste.

Reasoning

Strengths include a novel training paradigm using community feedback (citations) for preference modeling and alignment, with strong empirical results outperforming SOTA LLMs. Weaknesses include reliance on citation-based proxy for impact, which may not fully capture scientific taste, and limited scope to idea generation rather than full research automation.

Read-first score

Read-first score 51.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 60.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,experiment

Topical relevance 42%
50

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 62.

Keyword Scores

AI scientist
9
automated scientific discovery
8
AI for scientific research
8
scientific discovery agent
8
autonomous research agent
7
automated research
7
research automation
6
literature review agent
2
experiment design agent
2
survey generation
1
automated experimentation
1
paper writing agent
1

Deep Analysis

Innovations

  • Reinforcement Learning from Community Feedback (RLCF) paradigm for learning scientific taste from large-scale community signals
  • Formulation of scientific taste learning as a preference modeling and alignment problem
  • Scientific Judge model trained on 700K field- and time-matched high- vs. low-citation paper pairs to evaluate research ideas
  • Scientific Thinker policy model trained via preference alignment using Scientific Judge as a reward model to generate high-impact research ideas

Methodology

The authors propose RLCF, first training a Scientific Judge on 700K matched pairs of high- and low-citation papers to model scientific preference, then using it as a reward model to align a policy model, Scientific Thinker, to propose ideas with high potential impact.

Key Results

Scientific Judge outperforms state-of-the-art LLMs (GPT-5.2, Gemini 3 Pro) and generalizes to future years, unseen fields, and peer-review preferences; Scientific Thinker generates research ideas with higher potential impact than baselines.

Tags

CL