Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies

arXiv 2026 66.4 method

TLDR

Introduces TaxoBench to evaluate deep research agents on paper retrieval and taxonomy organization, revealing a dual bottleneck in capability and alignment.

Reasoning

The paper's strength lies in its novel benchmark (TaxoBench) with expert taxonomies and new metrics, providing a rigorous evaluation of both retrieval and hierarchical organization. Weaknesses include limited scope (only LLM surveys) and potential over-reliance on single expert references, though they attempt to partition findings into capability and alignment.

Read-first score

Read-first score 66.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 57.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,metric

Reproducibility 25%
73

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=github

Topical relevance 42%
47.5

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

FrontierMethodology anchorReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 87.

Keyword Scores

literature review agent
9
survey generation
9
automated research
7
research automation
7
autonomous research agent
6
AI for scientific research
5
automated scientific discovery
4
paper writing agent
3
scientific discovery agent
3
AI scientist
2
automated experimentation
1
experiment design agent
1

Deep Analysis

Innovations

  • TaxoBench: a benchmark of 72 LLM surveys with expert-authored taxonomy trees and paper-to-category mappings for evaluating retrieval and hierarchical organization
  • New hierarchical evaluation metrics: Unordered Semantic Tree Edit Distance (US-TED/US-NTED) and Semantic Path Similarity (Sem-Path) that capture taxonomy structure beyond flat clustering
  • Partitioning of evaluation into capability-based (reference-free) and alignment-based (reference-dependent) groups to separate model failure from disagreement with a single expert reference

Methodology

TaxoBench is constructed from 72 highly cited LLM surveys with expert taxonomies and 3,815 papers. It evaluates retrieval via Recall/Precision/F1 and organization at leaf and hierarchy levels using new metrics US-TED and Sem-Path. Two modes (Deep Research end-to-end and Bottom-Up organization-only) are tested on 7 Deep Research Agents and 12 frontier LLMs, with findings split into capability and alignment groups.

Key Results

The best agent retrieves only 20.92% of expert-cited papers; model taxonomies exhibit 75.9% sibling overlap, 51.2% MECE violations, and 83.4% structural imbalance (capability). All LLMs achieve Sem-Path 28-29%, far below human annotators' 47-58% (alignment).

Limitations

  • Reliance on a single expert-authored taxonomy as reference for alignment evaluation, though mitigated by partitioning findings into reference-free capability metrics

Tags

CL