ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
TLDR
Introduces Chain-of-Evidence framework and ScientistOne system to ensure verifiability in autonomous research, achieving zero hallucinated references and human-level performance.
Reasoning
The paper's strength lies in its novel verifiability framework and comprehensive empirical validation across multiple tasks, addressing a critical flaw in existing autonomous research agents. Weaknesses include limited discussion of scalability and potential reliance on specific task domains, though the abstract is strong overall.
Read-first score
Read-first score 62.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 86.
Field roles
Rank sensitivity
Stability: volatile; rank range: 84.
Keyword Scores
Deep Analysis
Innovations
- Chain-of-Evidence (CoE), a verifiability framework requiring every claim to be traceable to its evidence source
- ScientistOne, an end-to-end autonomous research system that maintains evidence chains by construction throughout literature review, solution discovery, and paper writing
- CoE Audit, a post-hoc audit with four integrity checks (score verification, specification violation, reference verification, method-code alignment) applicable uniformly to all systems
Methodology
The authors propose the Chain-of-Evidence framework and build ScientistOne, an autonomous research agent that enforces evidence chains across literature review, solution discovery, and writing. They evaluate using CoE Audit on 75 papers from five systems across five frontier research tasks, and test generalization on six additional tasks spanning medical imaging, fine-grained recognition, 3D perception, and language modeling.
Key Results
ScientistOne achieves zero hallucinated references (0/337), perfect score verification (12/12), and the highest method-code alignment (14/15), matching or exceeding human expert performance on all five tasks. It generalizes to six new tasks, reaching state-of-the-art on Parameter Golf and gold medals on MLE-Bench tasks where baselines fail entirely, while baselines exhibit systematic failures such as 21% hallucinated references and score verification as low as 42%.