Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys

arXiv 2025 48.6 method

TLDR

SurveyEval is a benchmark for evaluating LLM-generated academic surveys across quality, coherence, and accuracy.

Reasoning

The paper introduces a comprehensive evaluation benchmark with three dimensions and seven subjects, using human references to improve alignment. Its strength lies in addressing the underexplored evaluation challenge, but it does not propose new generation methods, limiting its scope to assessment.

Read-first score

Read-first score 48.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 47.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,result

Topical relevance 42%
39.2

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 16.

Keyword Scores

survey generation
8
literature review agent
7
AI for scientific research
6
automated research
5
paper writing agent
5
research automation
5
autonomous research agent
3
AI scientist
2
automated scientific discovery
2
scientific discovery agent
2
automated experimentation
1
experiment design agent
1

Deep Analysis

Innovations

  • SurveyEval benchmark evaluating surveys on overall quality, outline coherence, and reference accuracy
  • Augmentation of LLM-as-a-Judge with human references to improve evaluation-human alignment
  • Multi-subject evaluation across 7 subjects

Methodology

SurveyEval evaluates automatically generated surveys on three dimensions: overall quality, outline coherence, and reference accuracy. The benchmark covers 7 subjects and uses an LLM-as-a-Judge framework augmented with human references to align with human judgments.

Key Results

General long-text or paper-writing systems produce lower-quality surveys, while specialized survey-generation systems achieve substantially higher quality.

Tags

CLAI