Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?

arXiv 2025 68.6 method

TLDR

Evaluation of Sakana's AI Scientist reveals critical flaws: poor literature reviews, high experiment failure rate, minimal code changes, and hallucinated results.

Reasoning

The paper provides a thorough independent evaluation with concrete metrics (42% experiment failure, median 5 citations, etc.), highlighting both strengths (leap forward) and weaknesses. However, it focuses on a single system and does not propose new methods or solutions.

Read-first score

Read-first score 68.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 99.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Topical relevance 42%
82.5

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Methodology quality 25%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=evaluation,experiment,result

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code

Field roles

FrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 82.

Keyword Scores

AI scientist
10
automated scientific discovery
9
autonomous research agent
9
automated research
9
research automation
9
scientific discovery agent
9
automated experimentation
8
AI for scientific research
8
paper writing agent
8
literature review agent
7
experiment design agent
7
survey generation
6

Deep Analysis

Innovations

  • Introduces the term 'Artificial Research Intelligence (ARI)' to describe AI's ability to autonomously conduct research.
  • Provides the first thorough independent evaluation of Sakana's 'AI Scientist' system.

Methodology

The authors evaluated the AI Scientist by analyzing its literature reviews for novelty assessment, tracking experiment execution success rates and code modifications, and examining generated manuscripts for citation quality, structural errors, and hallucinated results. They also measured cost and human involvement time.

Key Results

42% of experiments failed due to coding errors, code changes averaged only 8% more characters, manuscripts had a median of 5 mostly outdated citations, frequent structural errors and hallucinated results, yet the system produced full papers for $6–15 with 3.5 hours of human involvement.

Tags

IRAILG