Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Survey

arXiv 2026 59.8 method

TLDR

Proposes DeepSurvey-Bench, a benchmark to evaluate academic value of automatically generated scientific surveys, addressing flaws in existing surface-level metrics.

Reasoning

The paper identifies a clear gap in evaluating survey generation quality beyond surface metrics, and introduces a multi-dimensional academic value criteria. However, the abstract is cut off, so full experimental validation is not visible, and the benchmark's novelty may be incremental.

Read-first score

Read-first score 59.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 49.

Methodology quality 25%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,benchmark,dataset,evaluation,experiment,metric,result

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Topical relevance 42%
40.8

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 78.

Keyword Scores

survey generation
10
literature review agent
7
AI for scientific research
6
research automation
6
automated research
5
paper writing agent
5
automated scientific discovery
3
AI scientist
2
scientific discovery agent
2
autonomous research agent
1
automated experimentation
1
experiment design agent
1

Deep Analysis

Innovations

  • Proposes DeepSurvey-Bench, a benchmark for evaluating academic value of automatically generated scientific surveys.
  • Defines a comprehensive academic value evaluation criteria with three dimensions: informational value, scholarly communication value, and research guidance value.
  • Constructs a reliable dataset with academic value annotations, moving beyond flawed selection criteria like citation counts and structural coherence.
  • Evaluates deep academic value (core research objectives, critical analysis) rather than surface-level metrics.

Methodology

The benchmark defines a three-dimensional evaluation criteria for academic value. A dataset is constructed with annotations based on these dimensions. Generated surveys are evaluated against this benchmark, and results are compared to human assessments.

Key Results

The benchmark demonstrates high consistency with human performance in assessing the academic value of generated surveys.

Tags

AICL