Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

When AI Does Science: Evaluating the Autonomous AI Scientist KOSMOS in Radiation Biology

arXiv 2025 64.7 system, benchmark, application

TLDR

Evaluates KOSMOS, an autonomous AI scientist, on three radiation biology hypotheses using random-gene null benchmarks, finding one valid discovery.

Reasoning

The paper provides a rigorous evaluation with null models and real datasets, but is limited to three hypotheses in a specific domain. Strengths include clear methodology and empirical results; weaknesses include narrow scope and mixed success.

Read-first score

Read-first score 64.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 72.

Methodology quality 25%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,benchmark,evaluation,metric,result

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Topical relevance 42%
60

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 21.

Keyword Scores

AI scientist
10
autonomous research agent
9
scientific discovery agent
9
automated scientific discovery
8
AI for scientific research
8
automated research
7
research automation
6
literature review agent
5
automated experimentation
4
experiment design agent
3
survey generation
2
paper writing agent
1

Deep Analysis

Innovations

  • Rigorous evaluation of an autonomous AI scientist (KOSMOS) using simple random-gene null benchmarks
  • Application of null-model auditing to AI-generated hypotheses in radiation biology

Methodology

KOSMOS, an autonomous AI scientist, generated three hypotheses for radiation biology problems. Each hypothesis was tested against random-gene null benchmarks using metrics such as Spearman correlation, empirical p-values, and concordance index to assess whether the AI's predictions outperformed random gene sets.

Key Results

One hypothesis (CDO1 predicting radiation-response modules) was strongly supported (r=0.70, empirical p=0.0039); a 12-gene signature for prostate radiotherapy outcome was significant but had a non-unique effect size (C-index=0.61, p=0.017); and a DDR-p53 hypothesis was not supported (rho=-0.40, p=0.76, indistinguishable from random).

Limitations

  • KOSMOS produced one false hypothesis and one uncertain result, demonstrating that AI-generated hypotheses can be misleading without rigorous auditing.
  • The significant 12-gene signature had a non-unique effect size, limiting its practical value.
  • Evaluation was restricted to three specific problems in radiation biology, so generalizability to other domains is unknown.

Tags

AICL