Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

arXiv 2026 61.5 method

TLDR

Introduces AutoResearchEval, evaluating 8 agent-harness combinations on 100 real-world research tasks, yielding 800 trajectories and a 45-pattern failure taxonomy centered on missing metacognitive loop.

Reasoning

The paper's strength is its large-scale, process-level diagnostic evaluation with artifact visibility across the full research lifecycle. However, the abstract lacks quantitative results and details on the validation of the agent-as-a-judge pipeline, limiting assessment of reliability.

Read-first score

Read-first score 61.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 68.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,benchmark,evaluation

Topical relevance 42%
56.7

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=artifact

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 20.

Keyword Scores

automated research
9
autonomous research agent
8
AI for scientific research
8
research automation
8
AI scientist
7
automated scientific discovery
6
automated experimentation
5
paper writing agent
5
scientific discovery agent
5
experiment design agent
4
literature review agent
2
survey generation
1

Deep Analysis

Innovations

  • AutoResearchEval: a benchmark of 100 real-world frontier research tasks across 7 scientific domains and the full research lifecycle, with process-level annotation.
  • AutoResearch Failure Taxonomy (ARFT): a framework of 45 empirically-grounded failure patterns derived from 800 agent trajectories.
  • Human-calibrated agent-as-a-judge pipeline for scalable fine-grained attribution of failures across trajectories and artifacts.
  • Identification of the lack of a metacognitive loop as the overarching limitation explaining diverse failure patterns in current autoresearch agents.

Methodology

They constructed AutoResearchEval with 100 tasks grounded in published frontier science, covering ideation, retrieval, execution, analysis, writing, and review. Eight harness-model combinations were evaluated, producing 800 agent trajectories with process-level annotations. A human-calibrated agent-as-a-judge pipeline inspected trajectories and intermediate artifacts to derive the failure taxonomy.

Key Results

Analysis of 800 trajectories revealed 45 failure patterns converging on a single overarching limitation: the absence of a metacognitive loop. These patterns consistently recurred across all 8 harness-model combinations, including the strongest models, indicating a model-level deficit rather than a scaffold-specific issue.

Limitations

  • Whether orchestration-level interventions can close the metacognitive loop gap is an open question not tested in this work.

Tags