Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

arXiv 2026 62.7 method

TLDR

Introduces Chain-of-Evidence framework and ScientistOne system to ensure verifiability in autonomous research, achieving zero hallucinated references and human-level performance.

Reasoning

The paper's strength lies in its novel verifiability framework and comprehensive empirical validation across multiple tasks, addressing a critical flaw in existing autonomous research agents. Weaknesses include limited discussion of scalability and potential reliance on specific task domains, though the abstract is strong overall.

Read-first score

Read-first score 62.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 86.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Topical relevance 42%
71.7

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,evaluation

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code

Field roles

FrontierBridge

Rank sensitivity

Stability: volatile; rank range: 84.

Keyword Scores

autonomous research agent
10
automated scientific discovery
9
automated research
9
research automation
9
scientific discovery agent
9
AI scientist
8
AI for scientific research
8
paper writing agent
8
literature review agent
7
survey generation
4
automated experimentation
3
experiment design agent
2

Deep Analysis

Innovations

  • Chain-of-Evidence (CoE), a verifiability framework requiring every claim to be traceable to its evidence source
  • ScientistOne, an end-to-end autonomous research system that maintains evidence chains by construction throughout literature review, solution discovery, and paper writing
  • CoE Audit, a post-hoc audit with four integrity checks (score verification, specification violation, reference verification, method-code alignment) applicable uniformly to all systems

Methodology

The authors propose the Chain-of-Evidence framework and build ScientistOne, an autonomous research agent that enforces evidence chains across literature review, solution discovery, and writing. They evaluate using CoE Audit on 75 papers from five systems across five frontier research tasks, and test generalization on six additional tasks spanning medical imaging, fine-grained recognition, 3D perception, and language modeling.

Key Results

ScientistOne achieves zero hallucinated references (0/337), perfect score verification (12/12), and the highest method-code alignment (14/15), matching or exceeding human expert performance on all five tasks. It generalizes to six new tasks, reaching state-of-the-art on Parameter Golf and gold medals on MLE-Bench tasks where baselines fail entirely, while baselines exhibit systematic failures such as 21% hallucinated references and score verification as low as 42%.

Tags

AICLMA