Benchmarking AI scientists for omics data driven biological discovery
TLDR
Introduces BAISBench, a benchmark evaluating AI scientists on real single-cell transcriptomic data for cell annotation and biological discovery.
Reasoning
The paper addresses a clear gap by providing a realistic, data-driven benchmark for AI scientists in biology, using real datasets and published conclusions. Its main limitation is the narrow focus on single-cell omics, which may not capture broader biological discovery tasks.
Read-first score
Read-first score 76.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 70.
Field roles
Rank sensitivity
Stability: volatile; rank range: 30.
Keyword Scores
Deep Analysis
Innovations
- Introduction of BAISBench, a benchmark for evaluating AI scientists on real single-cell transcriptomic datasets with two tasks: cell type annotation and scientific discovery via multiple-choice questions from published studies.
- Provision of a human performance baseline using six graduate-level bioinformaticians to contextualize AI scientist performance.
Methodology
BAISBench comprises 15 expert-labeled single-cell transcriptomic datasets for cell type annotation and 193 multiple-choice questions derived from 41 published single-cell studies. Several representative AI scientists were evaluated, and six graduate-level bioinformaticians completed the same tasks to establish a human baseline.
Key Results
Current AI scientists fall short of fully autonomous biological discovery but demonstrate substantial potential in supporting data-driven biological research.