Agentic AutoSurvey: Let LLMs Survey LLMs
TLDR
A multi-agent framework for automated survey generation, achieving superior synthesis quality on LLM research topics.
Reasoning
The paper presents a clear multi-agent architecture with specialized agents and quantitative results on real topics, demonstrating significant improvement over baselines. However, the evaluation is limited to LLM-related topics and may not generalize to other scientific domains.
Read-first score
Read-first score 61.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 62.
Field roles
Rank sensitivity
Stability: volatile; rank range: 37.
Keyword Scores
Deep Analysis
Innovations
- Multi-agent framework with four specialized agents (Paper Search Specialist, Topic Mining & Clustering, Academic Survey Writer, Quality Evaluator) for automated survey generation
- 12-dimension evaluation capturing organization, synthesis integration, and critical analysis beyond basic metrics
Methodology
The system employs four specialized agents working in concert to generate literature surveys. Experiments were conducted on six representative LLM research topics from COLM 2024 categories, processing 75–443 papers per topic (847 total) and comparing against the AutoSurvey baseline using a 12-dimension evaluation.
Key Results
The multi-agent approach achieved a score of 8.18/10 compared to AutoSurvey's 4.77/10, with high citation coverage (often ≥80% on 75–100-paper sets).
Limitations
- Lower citation coverage on very large paper sets (e.g., RLHF)
- Evaluation limited to six topics from COLM 2024 categories