Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

Benchmarking AI scientists for omics data driven biological discovery

arXiv 2025 76.8 method

TLDR

Introduces BAISBench, a benchmark evaluating AI scientists on real single-cell transcriptomic data for cell annotation and biological discovery.

Reasoning

The paper addresses a clear gap by providing a realistic, data-driven benchmark for AI scientists in biology, using real datasets and published conclusions. Its main limitation is the narrow focus on single-cell omics, which may not capture broader biological discovery tasks.

Read-first score

Read-first score 76.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 70.

Methodology quality 25%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,benchmark,dataset,evaluation,experiment,result

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Reproducibility 25%
81

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=dataset,github

Topical relevance 42%
58.3

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

FrontierBridgeMethodology anchorReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 30.

Keyword Scores

AI scientist
10
automated scientific discovery
9
AI for scientific research
9
scientific discovery agent
9
autonomous research agent
8
automated research
8
research automation
8
automated experimentation
3
literature review agent
2
experiment design agent
2
survey generation
1
paper writing agent
1

Deep Analysis

Innovations

  • Introduction of BAISBench, a benchmark for evaluating AI scientists on real single-cell transcriptomic datasets with two tasks: cell type annotation and scientific discovery via multiple-choice questions from published studies.
  • Provision of a human performance baseline using six graduate-level bioinformaticians to contextualize AI scientist performance.

Methodology

BAISBench comprises 15 expert-labeled single-cell transcriptomic datasets for cell type annotation and 193 multiple-choice questions derived from 41 published single-cell studies. Several representative AI scientists were evaluated, and six graduate-level bioinformaticians completed the same tasks to establish a human baseline.

Key Results

Current AI scientists fall short of fully autonomous biological discovery but demonstrate substantial potential in supporting data-driven biological research.

Tags

AIMAGN