Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

Kosmos: An AI Scientist for Autonomous Discovery

arXiv 2025 73.1 method

TLDR

Kosmos is an AI scientist that autonomously performs cycles of data analysis, literature search, and hypothesis generation over 12 hours, producing traceable scientific reports with high accuracy.

Reasoning

Strengths include a novel structured world model enabling coherent long-horizon research and empirical validation with human evaluators showing 79.4% accuracy and significant time savings. Weaknesses are its limitation to data-driven discovery, scalability only tested up to 20 cycles, and an incomplete abstract.

Read-first score

Read-first score 73.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 96.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Topical relevance 42%
80

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,dataset,result

Reproducibility 25%
50

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,dataset,reproduce

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 32.

Keyword Scores

AI scientist
10
automated scientific discovery
10
scientific discovery agent
10
autonomous research agent
9
automated research
9
AI for scientific research
9
research automation
9
paper writing agent
8
literature review agent
7
survey generation
6
automated experimentation
5
experiment design agent
4

Deep Analysis

Innovations

  • Structured world model that shares information between data analysis and literature search agents, enabling coherent long-horizon discovery over 200 agent rollouts.
  • Iterative cycles of parallel data analysis, literature search, and hypothesis generation, culminating in synthesized scientific reports.
  • Traceable reasoning via citation of all report statements with code or primary literature.
  • Linear scaling of valuable scientific findings with the number of cycles (tested up to 20 cycles).
  • Autonomous reproduction of unpublished findings and generation of novel contributions across metabolomics, materials science, neuroscience, and statistical genetics.

Methodology

Kosmos takes an open-ended objective and a dataset, then runs for up to 12 hours performing iterative cycles of parallel data analysis, literature search, and hypothesis generation. A structured world model enables information sharing between a data analysis agent and a literature search agent. Discoveries are synthesized into scientific reports where every statement is cited with code or primary literature.

Key Results

Independent scientists rated 79.4% of Kosmos report statements as accurate; a single 20-cycle run was judged equivalent to 6 months of researcher time on average, and the number of valuable findings scaled linearly with cycles. Seven discoveries spanned multiple fields, with three reproducing unpublished/preprinted results and four being novel contributions.

Limitations

  • Report statement accuracy is 79.4%, meaning 20.6% of statements may be inaccurate.
  • Scaling of findings was tested only up to 20 cycles; behavior beyond that is unknown.

Tags

AI