SurveyGen: Quality-Aware Scientific Survey Generation with Large Language Models
TLDR
SurveyGen introduces a large-scale dataset and quality-aware framework for evaluating and improving LLM-based automatic survey generation.
Reasoning
The paper's strength lies in its large, annotated dataset and novel quality-aware retrieval framework, enabling systematic evaluation. Weaknesses include the finding that fully automatic generation still underperforms in citation quality and critical analysis, limiting practical applicability.
Read-first score
Read-first score 59.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 52.
Field roles
Rank sensitivity
Stability: volatile; rank range: 65.
Keyword Scores
Deep Analysis
Innovations
- SurveyGen dataset: over 4,200 human-written surveys with 242,143 cited references and quality-related metadata
- QUAL-SG framework: quality-aware survey generation that enhances RAG by incorporating quality indicators into literature retrieval to select higher-quality source papers
Methodology
The authors construct the SurveyGen dataset of human-written surveys with quality metadata, then propose QUAL-SG, a quality-aware RAG framework that uses quality indicators to retrieve and select higher-quality papers. They systematically evaluate state-of-the-art LLMs under varying levels of human involvement, from fully automatic to human-guided generation.
Key Results
Semi-automatic pipelines can achieve partially competitive outcomes, but fully automatic survey generation still suffers from low citation quality and limited critical analysis.
Limitations
- Fully automatic survey generation still suffers from low citation quality and limited critical analysis.