ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure
TLDR
A benchmark evaluating LLMs' scientific hypothesis generation under progressive information disclosure, measuring innovativeness and grounded reasoning.
Reasoning
Strengths include a novel progressive disclosure framework and automated semantic evaluation. Weaknesses are reliance on semantic similarity, which may not capture true creativity, and the abstract cuts off, leaving results incomplete.
Read-first score
Read-first score 50, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.
Field roles
Rank sensitivity
Stability: volatile; rank range: 35.
Keyword Scores
Deep Analysis
Innovations
- Progressive information disclosure framework for evaluating hypothesis generation, from minimal context to full experimental details.
- Evaluation via automated semantic similarity of atomic claims to measure divergence from ground-truth conclusions.
- Benchmarking of multiple LLMs (GPT-5, GPT-5.4, Gemini 2.5 pro, Gemini 3.1 pro preview) on 45 recent scientific papers across materials science subfields.
Methodology
Models are given a research topic and question from a recent paper, with technical details progressively revealed. At each stage, they generate hypotheses, which are decomposed into atomic claims and compared to the original paper's conclusions using automated semantic similarity to quantify alignment and divergence.
Key Results
GPT-5.4 and Gemini 3.1 pro outperform their predecessors; GPT-5.4 maintains a 0.7 F1 score alignment with ground truth even under minimal context.