Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

AutoResearch: Insight In, Hallucination Out

arXiv 2026 60.5 method

TLDR

AutoResearch is a two-stage autonomous research system that generates grounded research ideas and executes experiments with evidence-based review, improving benchmark performance while reducing unreliable results.

Reasoning

The paper presents a clear two-stage architecture with mechanisms for grounded idea generation and evidence-based execution, supported by benchmark evaluations and audit comparisons. However, the abstract provides limited detail on baselines, scope, and limitations, making full assessment difficult.

Read-first score

Read-first score 60.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 71.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,experiment,metric,result

Topical relevance 42%
59.2

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 26.

Keyword Scores

autonomous research agent
9
automated research
9
automated scientific discovery
8
automated experimentation
8
AI for scientific research
8
research automation
8
scientific discovery agent
8
experiment design agent
7
AI scientist
6
literature review agent
0
survey generation
0
paper writing agent
0

Deep Analysis

Innovations

  • Two-stage system connecting Idea Generation and Idea Execution to ensure scientific grounding
  • Idea Generation: continuous integration of emerging research signals with domain knowledge, identification of transferable mechanistic insights, multi-model generation and cross-review to produce grounded, testable research plans
  • Idea Execution: coordinated agents that decompose plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting conclusions
  • Evidence-conditioned decision-making to continue, revise, or terminate research directions, and detection/correction of unreliable experimental results

Methodology

AutoResearch is a two-stage system where Idea Generation uses multi-model generation and cross-review with integrated research signals and domain knowledge to produce testable plans, and Idea Execution deploys coordinated agents that decompose plans, iteratively implement and diagnose experiments, and apply independent evidence-based review. The system is evaluated on cross-modal retrieval, systems optimization, and benchmark-driven ML, using metrics like mean Recall and audit-confirmed issue events.

Key Results

On the RSICD benchmark, an AutoResearch-generated idea improved mean Recall from 32.84 to 34.69, with only 5 audit-confirmed issue events compared to 11-27 for other autonomous systems, demonstrating measurable progress, error detection, and evidence-conditioned decisions.

Tags