SurveyLens: A Discipline-Aware Benchmark for Automatic Survey Generation
TLDR
SurveyLens is the first discipline-aware benchmark for automatic survey generation, evaluating 11 systems across 10 disciplines.
Reasoning
The paper introduces a novel benchmark with a curated dataset and dual-lens evaluation framework, addressing a gap in cross-discipline ASG assessment. Its strengths include comprehensive system evaluation and practical guidance, but it is limited to survey generation rather than broader scientific discovery.
Read-first score
Read-first score 54.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 55.
Field roles
Rank sensitivity
Stability: volatile; rank range: 22.
Keyword Scores
Deep Analysis
Innovations
- First discipline-aware benchmark for Automatic Survey Generation (SurveyLens)
- SurveyLens-1k: a curated dataset of 1,000 human-written surveys across 10 disciplines
- Dual-lens evaluation framework combining discipline-aware rubric scoring and reference-based alignment to human-written surveys
Methodology
The paper introduces SurveyLens, a benchmark with 1,000 human-written surveys from 10 disciplines, and evaluates 11 systems (vanilla LLMs, ASG systems, Deep Research agents) using a dual-lens approach: discipline-specific rubric scoring and alignment with reference surveys.
Key Results
Deep Research agents are the only paradigm robust across all 10 disciplines, ASG systems excel at structural planning, and all paradigms show weakness in reference quality.