AI scientists produce results without reasoning scientifically
TLDR
LLM-based scientific agents execute workflows but fail to exhibit epistemic reasoning, ignoring evidence in 68% of traces.
Reasoning
The paper's strength lies in its large-scale empirical evaluation (25,000+ runs) across eight domains, systematically decomposing base model vs. scaffold contributions. However, it focuses narrowly on LLM-based agents and may not generalize to other AI scientist paradigms; the abstract lacks details on specific domains or error analysis.
Read-first score
Read-first score 59.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 67.
Field roles
Rank sensitivity
Stability: volatile; rank range: 25.
Keyword Scores
Deep Analysis
Innovations
- Systematic decomposition of performance and behavior into base model versus agent scaffold contributions
- Epistemological analysis of agent reasoning traces across workflow execution and hypothesis-driven inquiry
- Demonstration that outcome-based evaluation masks failures in scientific reasoning patterns
Methodology
Over 25,000 agent runs across eight scientific domains were analyzed through two lenses: a variance decomposition of base model and scaffold contributions to performance, and a behavioral analysis of reasoning traces for evidence use, refutation-driven belief revision, and multi-test convergence.
Key Results
The base model explains 41.4% of performance variance versus 1.5% for the scaffold; evidence is ignored in 68% of traces, refutation-driven belief revision occurs in only 26%, and convergent multi-test evidence is rare, with these patterns persisting across task types and compounding unreliability over repeated trials.
Limitations
- Outcome-based evaluation cannot detect epistemic failures in agent reasoning
- Scaffold engineering alone cannot repair the lack of scientific reasoning patterns
- Unreliability compounds across repeated trials in epistemically demanding domains