Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

Generating Literature-Driven Scientific Theories at Scale

arXiv 2026 60.5 method

TLDR

Automated theory generation from scientific literature at scale, using 13.7k papers to synthesize 2.9k theories, outperforming parametric methods.

Reasoning

The paper presents a novel approach to automated scientific discovery by focusing on theory building rather than experiment generation, using large-scale literature grounding. Strengths include empirical validation with real papers and future predictions; weaknesses may include limited discussion of theory quality metrics and potential biases in literature selection.

Read-first score

Read-first score 60.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 71.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=evaluation,experiment,metric,result

Topical relevance 42%
59.2

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 26.

Keyword Scores

automated scientific discovery
10
AI scientist
9
AI for scientific research
9
scientific discovery agent
9
automated research
8
research automation
8
autonomous research agent
6
literature review agent
5
paper writing agent
3
survey generation
2
automated experimentation
1
experiment design agent
1

Deep Analysis

Innovations

  • Formulation of theory synthesis from scientific literature as a novel problem in automated scientific discovery
  • Large-scale study generating 2.9k theories from 13.7k source papers
  • Comparison of literature-grounding versus parametric LLM knowledge for theory generation
  • Evaluation of accuracy-focused versus novelty-focused generation objectives

Methodology

The authors synthesize theories from a corpus of 13.7k scientific papers, generating 2.9k theories. They compare generation using literature-grounding (retrieval from papers) against parametric LLM memory, and examine accuracy-focused versus novelty-focused objectives. Evaluation measures how well theories match existing evidence and predict results from 4.6k subsequently-written papers.

Key Results

Literature-supported theory generation significantly outperforms parametric LLM memory in matching existing evidence and predicting future results from later papers.

Tags

CLAI