Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

Agentic AutoSurvey: Let LLMs Survey LLMs

arXiv 2025 61.3 method

TLDR

A multi-agent framework for automated survey generation, achieving superior synthesis quality on LLM research topics.

Reasoning

The paper presents a clear multi-agent architecture with specialized agents and quantitative results on real topics, demonstrating significant improvement over baselines. However, the evaluation is limited to LLM-related topics and may not generalize to other scientific domains.

Read-first score

Read-first score 61.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 62.

Methodology quality 25%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,baseline,evaluation,experiment,metric

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Topical relevance 42%
51.7

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 37.

Keyword Scores

survey generation
10
literature review agent
9
paper writing agent
8
research automation
7
automated research
6
AI for scientific research
6
autonomous research agent
5
automated scientific discovery
4
scientific discovery agent
3
AI scientist
2
automated experimentation
1
experiment design agent
1

Deep Analysis

Innovations

  • Multi-agent framework with four specialized agents (Paper Search Specialist, Topic Mining & Clustering, Academic Survey Writer, Quality Evaluator) for automated survey generation
  • 12-dimension evaluation capturing organization, synthesis integration, and critical analysis beyond basic metrics

Methodology

The system employs four specialized agents working in concert to generate literature surveys. Experiments were conducted on six representative LLM research topics from COLM 2024 categories, processing 75–443 papers per topic (847 total) and comparing against the AutoSurvey baseline using a 12-dimension evaluation.

Key Results

The multi-agent approach achieved a score of 8.18/10 compared to AutoSurvey's 4.77/10, with high citation coverage (often ≥80% on 75–100-paper sets).

Limitations

  • Lower citation coverage on very large paper sets (e.g., RLHF)
  • Evaluation limited to six topics from COLM 2024 categories

Tags

IRCLHC