Awesome Auto Research Hub 论文 · 数据集 · 项目
← 返回论文列表

How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

arXiv 2026 61.5 method

TLDR

Introduces AutoResearchEval, evaluating 8 agent-harness combinations on 100 real-world research tasks, yielding 800 trajectories and a 45-pattern failure taxonomy centered on missing metacognitive loop.

评分理由

The paper's strength is its large-scale, process-level diagnostic evaluation with artifact visibility across the full research lifecycle. However, the abstract lacks quantitative results and details on the validation of the agent-as-a-judge pipeline, limiting assessment of reliability.

Read-first 评分解释

综合优先阅读分 61.5,由主题、引用、图谱、方法、可复现性和近期性等信号加权得到。 原始总分保留为 68。

近期性 8%
100

使用温和的时间衰减,让近期论文更容易浮现,同时保留较早基础工作的价值。 年份:2026

方法质量 25%
80

检查可见的摘要与分析字段,寻找实验、数据集、基线、指标和局限性等方法证据。 命中信号:分析、基准、评估

主题相关性 42%
56.7

使用现有 LLM 关键词相关性评分,并归一化到 0-100。 关键词:AI scientist、automated scientific discovery、autonomous research agent、automated research、literature review agent、survey generation、automated experimentation、experiment design agent、AI for scientific research、paper writing agent、research automation、scientific discovery agent

可复现性 25%
38

检查链接和可见文本中的论文、代码、数据集、工件与仓库信号。 论文:有;代码:无;数据:无;命中信号:工件

研究版图角色

前沿论文方法锚点

排序敏感性

稳定性:volatile;排名波动范围:20。

关键词评分

automated research
9
autonomous research agent
8
AI for scientific research
8
research automation
8
AI scientist
7
automated scientific discovery
6
automated experimentation
5
paper writing agent
5
scientific discovery agent
5
experiment design agent
4
literature review agent
2
survey generation
1

深度分析

创新点

  • AutoResearchEval: a benchmark of 100 real-world frontier research tasks across 7 scientific domains and the full research lifecycle, with process-level annotation.
  • AutoResearch Failure Taxonomy (ARFT): a framework of 45 empirically-grounded failure patterns derived from 800 agent trajectories.
  • Human-calibrated agent-as-a-judge pipeline for scalable fine-grained attribution of failures across trajectories and artifacts.
  • Identification of the lack of a metacognitive loop as the overarching limitation explaining diverse failure patterns in current autoresearch agents.

方法

They constructed AutoResearchEval with 100 tasks grounded in published frontier science, covering ideation, retrieval, execution, analysis, writing, and review. Eight harness-model combinations were evaluated, producing 800 agent trajectories with process-level annotations. A human-calibrated agent-as-a-judge pipeline inspected trajectories and intermediate artifacts to derive the failure taxonomy.

关键结果

Analysis of 800 trajectories revealed 45 failure patterns converging on a single overarching limitation: the absence of a metacognitive loop. These patterns consistently recurred across all 8 harness-model combinations, including the strongest models, indicating a model-level deficit rather than a scaffold-specific issue.

局限性

  • Whether orchestration-level interventions can close the metacognitive loop gap is an open question not tested in this work.

标签