Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

The More You Automate, the Less You See: Hidden Pitfalls of AI Scientist Systems

arXiv 2025 67.9 method

TLDR

Identifies four failure modes in AI scientist systems and demonstrates them via controlled experiments on two open-source systems.

Reasoning

The paper's strength lies in its systematic identification and empirical demonstration of critical pitfalls (benchmark selection, data leakage, metric misuse, post-hoc bias) that undermine trust in AI-generated research. A weakness is that the abstract does not detail the severity or generalizability of the failures beyond the two tested systems, and the proposed solution (mandating trace logs) is a recommendation rather than a validated method.

Read-first score

Read-first score 67.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 84.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,experiment,metric

Topical relevance 42%
70

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=artifact,code

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 34.

Keyword Scores

AI scientist
10
automated scientific discovery
9
autonomous research agent
9
scientific discovery agent
9
automated research
8
AI for scientific research
8
research automation
8
automated experimentation
7
paper writing agent
7
experiment design agent
6
literature review agent
2
survey generation
1

Deep Analysis

Innovations

  • Identification of four failure modes in AI scientist systems: inappropriate benchmark selection, data leakage, metric misuse, and post-hoc selection bias.
  • Controlled experimental design to isolate each failure mode while addressing evaluation challenges unique to AI scientist systems.
  • Demonstration that access to trace logs and code from the automated workflow enables far more effective failure detection than examining the final paper alone.
  • Recommendation that journals and conferences mandate submission of trace logs and code alongside AI-generated papers for transparency, accountability, and reproducibility.

Methodology

The authors design controlled experiments that isolate each of the four failure modes, assess two prominent open-source AI scientist systems, and compare failure detection using trace logs and code versus the final paper alone.

Key Results

The assessment reveals several failures across a spectrum of severity that can be easily overlooked in practice; trace logs and code enable far more effective detection of these failures than examining the final paper alone.

Tags

AIDL