Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation

arXiv 2025 58.8 method

TLDR

Introduces SurGE, a benchmark and evaluation framework for scientific survey generation, revealing performance gaps in LLM-based methods.

Reasoning

The paper addresses a critical gap by providing a standardized benchmark and evaluation framework for survey generation, with real data and open-source resources. However, it is limited to computer science and does not propose new methods, only evaluating existing ones.

Read-first score

Read-first score 58.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 47.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Reproducibility 25%
81

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=code,github

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation

Topical relevance 42%
39.2

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

FrontierBridgeReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 81.

Keyword Scores

survey generation
10
literature review agent
9
paper writing agent
7
AI for scientific research
5
automated research
4
research automation
4
AI scientist
3
automated scientific discovery
2
autonomous research agent
2
scientific discovery agent
1
automated experimentation
0
experiment design agent
0

Deep Analysis

Innovations

  • Introduction of SurGE, a benchmark for scientific survey generation in computer science, including test instances with expert-written surveys and cited references, and a large-scale academic corpus.
  • Proposal of an automated evaluation framework that measures survey quality across four dimensions: comprehensiveness, citation accuracy, structural organization, and content quality.

Methodology

We construct SurGE with test instances comprising topic descriptions, expert-written surveys, and their full set of cited references, alongside a corpus of over one million papers. We then evaluate diverse LLM-based methods using an automated framework that scores generated surveys on comprehensiveness, citation accuracy, structural organization, and content quality.

Key Results

Evaluation reveals a significant performance gap, with even advanced agentic frameworks struggling with the complexities of survey generation.

Tags

CLAIIR