Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

SurveyLens: A Discipline-Aware Benchmark for Automatic Survey Generation

arXiv 2026 54.4 method

TLDR

SurveyLens is the first discipline-aware benchmark for automatic survey generation, evaluating 11 systems across 10 disciplines.

Reasoning

The paper introduces a novel benchmark with a curated dataset and dual-lens evaluation framework, addressing a gap in cross-discipline ASG assessment. Its strengths include comprehensive system evaluation and practical guidance, but it is limited to survey generation rather than broader scientific discovery.

Read-first score

Read-first score 54.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 55.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,dataset,evaluation

Topical relevance 42%
45.8

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 22.

Keyword Scores

survey generation
10
literature review agent
9
paper writing agent
7
autonomous research agent
6
automated research
5
research automation
5
AI for scientific research
4
automated scientific discovery
3
AI scientist
2
scientific discovery agent
2
automated experimentation
1
experiment design agent
1

Deep Analysis

Innovations

  • First discipline-aware benchmark for Automatic Survey Generation (SurveyLens)
  • SurveyLens-1k: a curated dataset of 1,000 human-written surveys across 10 disciplines
  • Dual-lens evaluation framework combining discipline-aware rubric scoring and reference-based alignment to human-written surveys

Methodology

The paper introduces SurveyLens, a benchmark with 1,000 human-written surveys from 10 disciplines, and evaluates 11 systems (vanilla LLMs, ASG systems, Deep Research agents) using a dual-lens approach: discipline-specific rubric scoring and alignment with reference surveys.

Key Results

Deep Research agents are the only paradigm robust across all 10 disciplines, ASG systems excel at structural planning, and all paradigms show weakness in reference quality.

Tags

CL