Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies
TLDR
Introduces TaxoBench to evaluate deep research agents on paper retrieval and taxonomy organization, revealing a dual bottleneck in capability and alignment.
Reasoning
The paper's strength lies in its novel benchmark (TaxoBench) with expert taxonomies and new metrics, providing a rigorous evaluation of both retrieval and hierarchical organization. Weaknesses include limited scope (only LLM surveys) and potential over-reliance on single expert references, though they attempt to partition findings into capability and alignment.
Read-first score
Read-first score 66.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 57.
Field roles
Rank sensitivity
Stability: volatile; rank range: 87.
Keyword Scores
Deep Analysis
Innovations
- TaxoBench: a benchmark of 72 LLM surveys with expert-authored taxonomy trees and paper-to-category mappings for evaluating retrieval and hierarchical organization
- New hierarchical evaluation metrics: Unordered Semantic Tree Edit Distance (US-TED/US-NTED) and Semantic Path Similarity (Sem-Path) that capture taxonomy structure beyond flat clustering
- Partitioning of evaluation into capability-based (reference-free) and alignment-based (reference-dependent) groups to separate model failure from disagreement with a single expert reference
Methodology
TaxoBench is constructed from 72 highly cited LLM surveys with expert taxonomies and 3,815 papers. It evaluates retrieval via Recall/Precision/F1 and organization at leaf and hierarchy levels using new metrics US-TED and Sem-Path. Two modes (Deep Research end-to-end and Bottom-Up organization-only) are tested on 7 Deep Research Agents and 12 frontier LLMs, with findings split into capability and alignment groups.
Key Results
The best agent retrieves only 20.92% of expert-cited papers; model taxonomies exhibit 75.9% sibling overlap, 51.2% MECE violations, and 83.4% structural imbalance (capability). All LLMs achieve Sem-Path 28-29%, far below human annotators' 47-58% (alignment).
Limitations
- Reliance on a single expert-authored taxonomy as reference for alignment evaluation, though mitigated by partitioning findings into reference-free capability metrics