Kosmos: An AI Scientist for Autonomous Discovery
TLDR
Kosmos is an AI scientist that autonomously performs cycles of data analysis, literature search, and hypothesis generation over 12 hours, producing traceable scientific reports with high accuracy.
Reasoning
Strengths include a novel structured world model enabling coherent long-horizon research and empirical validation with human evaluators showing 79.4% accuracy and significant time savings. Weaknesses are its limitation to data-driven discovery, scalability only tested up to 20 cycles, and an incomplete abstract.
Read-first score
Read-first score 73.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 96.
Field roles
Rank sensitivity
Stability: volatile; rank range: 32.
Keyword Scores
Deep Analysis
Innovations
- Structured world model that shares information between data analysis and literature search agents, enabling coherent long-horizon discovery over 200 agent rollouts.
- Iterative cycles of parallel data analysis, literature search, and hypothesis generation, culminating in synthesized scientific reports.
- Traceable reasoning via citation of all report statements with code or primary literature.
- Linear scaling of valuable scientific findings with the number of cycles (tested up to 20 cycles).
- Autonomous reproduction of unpublished findings and generation of novel contributions across metabolomics, materials science, neuroscience, and statistical genetics.
Methodology
Kosmos takes an open-ended objective and a dataset, then runs for up to 12 hours performing iterative cycles of parallel data analysis, literature search, and hypothesis generation. A structured world model enables information sharing between a data analysis agent and a literature search agent. Discoveries are synthesized into scientific reports where every statement is cited with code or primary literature.
Key Results
Independent scientists rated 79.4% of Kosmos report statements as accurate; a single 20-cycle run was judged equivalent to 6 months of researcher time on average, and the number of valuable findings scaled linearly with cycles. Seven discoveries spanned multiple fields, with three reproducing unpublished/preprinted results and four being novel contributions.
Limitations
- Report statement accuracy is 79.4%, meaning 20.6% of statements may be inaccurate.
- Scaling of findings was tested only up to 20 cycles; behavior beyond that is unknown.