Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
TLDR
SDABench evaluates LLMs on six scientific data analysis capabilities across five domains, finding strengths in descriptive analysis but weaknesses in assumption selection and mechanistic reasoning.
Reasoning
The paper introduces a well-structured benchmark with real and synthetic data, and a five-stage error analysis framework, which are strengths. However, it focuses narrowly on data analysis capabilities rather than full scientific discovery, and does not address autonomous research agents or experimentation.
Read-first score
Read-first score 52.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 50.
Field roles
Rank sensitivity
Stability: volatile; rank range: 39.
Keyword Scores
Deep Analysis
Innovations
- Capability-oriented evaluation framework reorganizing scientific analysis around six distinct claim types (descriptive, exploratory, inferential, predictive, causal, mechanistic) across five domains.
- Large-scale benchmark with both real (527) and synthetic (6000) data instances in multiple-choice and open-ended formats, constructed via an automated pipeline.
- Five-stage error analysis framework that pinpoints specific failure stages in LLM scientific reasoning (scope/variable identification, analytical procedure selection, variable relationship modeling, conclusion drawing).
Methodology
SDABench evaluates LLMs on six scientific analysis capabilities (descriptive, exploratory, inferential, predictive, causal, mechanistic) across five domains (Biology, Chemistry, Environment, Geography, Physics) using 527 real-data and 6000 synthetic instances. Both multiple-choice and open-ended formats assess 15 representative LLMs, and a five-stage error analysis framework is applied to identify where models fail in the reasoning process.
Key Results
LLMs perform well on descriptive analysis but degrade sharply on tasks requiring assumption selection, latent-process modeling, or mechanistic reasoning; more advanced models identify scope and variables but struggle with selecting analytical procedures, modeling variable relationships, and drawing valid conclusions.