SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
TLDR
SurveyEval is a benchmark for evaluating LLM-generated academic surveys across quality, coherence, and accuracy.
Reasoning
The paper introduces a comprehensive evaluation benchmark with three dimensions and seven subjects, using human references to improve alignment. Its strength lies in addressing the underexplored evaluation challenge, but it does not propose new generation methods, limiting its scope to assessment.
Read-first score
Read-first score 48.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 47.
Field roles
Rank sensitivity
Stability: volatile; rank range: 16.
Keyword Scores
Deep Analysis
Innovations
- SurveyEval benchmark evaluating surveys on overall quality, outline coherence, and reference accuracy
- Augmentation of LLM-as-a-Judge with human references to improve evaluation-human alignment
- Multi-subject evaluation across 7 subjects
Methodology
SurveyEval evaluates automatically generated surveys on three dimensions: overall quality, outline coherence, and reference accuracy. The benchmark covers 7 subjects and uses an LLM-as-a-Judge framework augmented with human references to align with human judgments.
Key Results
General long-text or paper-writing systems produce lower-quality surveys, while specialized survey-generation systems achieve substantially higher quality.