Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure

arXiv 2026 50 method

TLDR

A benchmark evaluating LLMs' scientific hypothesis generation under progressive information disclosure, measuring innovativeness and grounded reasoning.

Reasoning

Strengths include a novel progressive disclosure framework and automated semantic evaluation. Weaknesses are reliance on semantic similarity, which may not capture true creativity, and the abstract cuts off, leaving results incomplete.

Read-first score

Read-first score 50, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,experiment

Topical relevance 42%
40

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 35.

Keyword Scores

AI for scientific research
8
AI scientist
7
scientific discovery agent
7
automated scientific discovery
6
autonomous research agent
5
automated research
4
research automation
3
literature review agent
2
automated experimentation
2
experiment design agent
2
survey generation
1
paper writing agent
1

Deep Analysis

Innovations

  • Progressive information disclosure framework for evaluating hypothesis generation, from minimal context to full experimental details.
  • Evaluation via automated semantic similarity of atomic claims to measure divergence from ground-truth conclusions.
  • Benchmarking of multiple LLMs (GPT-5, GPT-5.4, Gemini 2.5 pro, Gemini 3.1 pro preview) on 45 recent scientific papers across materials science subfields.

Methodology

Models are given a research topic and question from a recent paper, with technical details progressively revealed. At each stage, they generate hypotheses, which are decomposed into atomic claims and compared to the original paper's conclusions using automated semantic similarity to quantify alignment and divergence.

Key Results

GPT-5.4 and Gemini 3.1 pro outperform their predecessors; GPT-5.4 maintains a 0.7 F1 score alignment with ground truth even under minimal context.

Tags

AI