Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

SurveyGen: Quality-Aware Scientific Survey Generation with Large Language Models

arXiv 2025 59.8 method

TLDR

SurveyGen introduces a large-scale dataset and quality-aware framework for evaluating and improving LLM-based automatic survey generation.

Reasoning

The paper's strength lies in its large, annotated dataset and novel quality-aware retrieval framework, enabling systematic evaluation. Weaknesses include the finding that fully automatic generation still underperforms in citation quality and critical analysis, limiting practical applicability.

Read-first score

Read-first score 59.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 52.

Methodology quality 25%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,dataset,evaluation,experiment,result

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Topical relevance 42%
43.3

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Field roles

FrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 65.

Keyword Scores

survey generation
10
automated research
7
research automation
7
literature review agent
6
AI for scientific research
6
paper writing agent
5
automated scientific discovery
4
AI scientist
2
scientific discovery agent
2
autonomous research agent
1
automated experimentation
1
experiment design agent
1

Deep Analysis

Innovations

  • SurveyGen dataset: over 4,200 human-written surveys with 242,143 cited references and quality-related metadata
  • QUAL-SG framework: quality-aware survey generation that enhances RAG by incorporating quality indicators into literature retrieval to select higher-quality source papers

Methodology

The authors construct the SurveyGen dataset of human-written surveys with quality metadata, then propose QUAL-SG, a quality-aware RAG framework that uses quality indicators to retrieve and select higher-quality papers. They systematically evaluate state-of-the-art LLMs under varying levels of human involvement, from fully automatic to human-guided generation.

Key Results

Semi-automatic pipelines can achieve partially competitive outcomes, but fully automatic survey generation still suffers from low citation quality and limited critical analysis.

Limitations

  • Fully automatic survey generation still suffers from low citation quality and limited critical analysis.

Tags

CLDLIR