SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
TLDR
Introduces SurGE, a benchmark and evaluation framework for scientific survey generation, revealing performance gaps in LLM-based methods.
Reasoning
The paper addresses a critical gap by providing a standardized benchmark and evaluation framework for survey generation, with real data and open-source resources. However, it is limited to computer science and does not propose new methods, only evaluating existing ones.
Read-first score
Read-first score 58.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 47.
Field roles
Rank sensitivity
Stability: volatile; rank range: 81.
Keyword Scores
Deep Analysis
Innovations
- Introduction of SurGE, a benchmark for scientific survey generation in computer science, including test instances with expert-written surveys and cited references, and a large-scale academic corpus.
- Proposal of an automated evaluation framework that measures survey quality across four dimensions: comprehensiveness, citation accuracy, structural organization, and content quality.
Methodology
We construct SurGE with test instances comprising topic descriptions, expert-written surveys, and their full set of cited references, alongside a corpus of over one million papers. We then evaluate diverse LLM-based methods using an automated framework that scores generated surveys on comprehensiveness, citation accuracy, structural organization, and content quality.
Key Results
Evaluation reveals a significant performance gap, with even advanced agentic frameworks struggling with the complexities of survey generation.