SurveyForge: On the Outline Heuristics, Memory-Driven Generation, and Multi-dimensional Evaluation for Automated Survey Writing
TLDR
SurveyForge improves automated survey writing by using outline heuristics, memory-driven generation, and multi-dimensional evaluation, outperforming prior methods.
Reasoning
The paper introduces a novel approach to address quality gaps in LLM-generated surveys, with strengths in outline analysis and citation accuracy. However, it focuses narrowly on survey writing rather than broader scientific discovery, and evaluation relies on win-rate comparisons without extensive real-world deployment.
Read-first score
Read-first score 50.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 53.
Field roles
Rank sensitivity
Stability: volatile; rank range: 31.
Keyword Scores
Deep Analysis
Innovations
- Outline generation heuristics based on logical structure of human-written outlines and retrieved domain articles
- Memory-driven generation with a scholar navigation agent that retrieves high-quality papers for content generation and refinement
- SurveyBench: a multi-dimensional evaluation benchmark with 100 human-written surveys for win-rate comparison across reference, outline, and content quality
Methodology
SurveyForge first generates an outline by analyzing the logical structure of human-written outlines and incorporating retrieved domain-related articles. Then, a scholar navigation agent retrieves high-quality papers from memory to automatically generate and refine the survey content. Evaluation is performed on SurveyBench, a benchmark of 100 human-written surveys, using win-rate comparison across three dimensions: reference, outline, and content quality.
Key Results
SurveyForge outperforms previous works such as AutoSurvey on the SurveyBench benchmark.