Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?

arXiv 2026 56 method

TLDR

Introduces SoundnessBench to test LLMs' ability to judge research idea soundness, finding pervasive optimism bias and unreliability.

Reasoning

Strengths include a novel, carefully curated benchmark with real ICLR submissions and multiple controls. Weaknesses are the narrow focus on ML proposals and lack of evaluation on full research pipeline automation.

Read-first score

Read-first score 56, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 45.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,experiment,result

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=artifact

Topical relevance 42%
37.5

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 77.

Keyword Scores

AI for scientific research
8
autonomous research agent
7
AI scientist
6
automated scientific discovery
5
automated research
5
scientific discovery agent
5
research automation
4
literature review agent
1
survey generation
1
automated experimentation
1
experiment design agent
1
paper writing agent
1

Deep Analysis

Innovations

  • Introduction of SoundnessBench, a curated benchmark of 1,099 machine-learning research proposals reconstructed from ICLR submissions with reviewer soundness sub-scores and source-paper audits.
  • Discovery of a pervasive optimism bias in frontier LLMs, where low-soundness proposals are frequently rated as sound under standard prompting.
  • Demonstration that aggressive prompting shifts errors from false positives to false negatives without improving overall evaluation reliability.
  • Rigorous control experiments for public-corpus contamination, paper-identifying phrases, surface features, and human audit quality to rule out single-confounder explanations.

Methodology

SoundnessBench was built from 1,099 ICLR submissions, reconstructing research proposals and labeling them with reviewer soundness sub-scores, then audited against source papers. Twelve frontier LLMs were evaluated under standard and aggressive prompting, measuring optimism bias and error-type shifts. Additional control experiments tested for contamination, surface features, and audit quality confounders.

Key Results

LLMs show a strong optimism bias, often rating unsound proposals as sound; aggressive prompting reduces false positives but increases false negatives, and no single confounder explains the behavior, indicating LLMs are unreliable as standalone first-gate scientific rigor evaluators.

Limitations

  • SoundnessBench measures recoverable proposal-stage soundness, not exact prediction of full-paper review outcomes.
  • The benchmark is limited to machine learning research proposals from ICLR, so findings may not generalize to other scientific domains.
  • Aggressive prompting only shifts error types (false positives to false negatives) without improving overall reliability.
  • Reconstructed proposals from submissions may introduce artifacts not present in original idea-stage proposals.

Tags

LG