Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research

arXiv 2025 52.3 method

TLDR

SPOT benchmark evaluates LLMs as verifiers of scientific manuscripts, finding low recall and precision, highlighting gaps in automated verification.

Reasoning

The paper introduces a novel benchmark (SPOT) for automated verification of scientific manuscripts using real published papers with known errors, which is a strength. However, the dataset is small (83 papers) and the evaluation shows very low performance, but the paper does not propose improvements or solutions, limiting its impact.

Read-first score

Read-first score 52.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 39.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,benchmark,dataset,metric

Reproducibility 25%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,dataset

Topical relevance 42%
32.5

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 76.

Keyword Scores

AI scientist
7
automated scientific discovery
6
AI for scientific research
6
automated research
5
research automation
5
autonomous research agent
4
scientific discovery agent
4
literature review agent
1
paper writing agent
1
survey generation
0
automated experimentation
0
experiment design agent
0

Deep Analysis

Innovations

  • Introduction of SPOT, a dataset of 83 published papers with 91 errors significant enough to prompt errata or retraction, cross-validated with authors and human annotators.
  • Proposing the use of LLMs as verifiers for automated academic verification of scientific manuscripts, a complementary role to generative AI co-scientists.

Methodology

State-of-the-art LLMs were evaluated on the SPOT dataset using recall and precision metrics, along with confidence estimates and consistency across eight independent runs. Qualitative analysis with domain experts was conducted to categorize model mistakes.

Key Results

No LLM exceeded 21.1% recall or 6.1% precision; o3 achieved the best scores while others were near zero. Confidence estimates were uniformly low, and models rarely rediscovered the same errors across runs.

Tags

CL