SGSimEval: A Comprehensive Multifaceted and Similarity-Enhanced Benchmark for Automatic Survey Generation Systems
TLDR
SGSimEval is a benchmark for evaluating automatic survey generation systems using multifaceted metrics including human preference and similarity.
Reasoning
The paper addresses limitations in existing evaluation methods for automatic survey generation by proposing a comprehensive benchmark that integrates outline, content, and reference assessments with LLM-based and quantitative metrics. Its strength lies in introducing human preference metrics and demonstrating consistency with human judgments, but it is narrowly focused on survey generation rather than broader scientific discovery.
Read-first score
Read-first score 51.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.
Field roles
Rank sensitivity
Stability: volatile; rank range: 24.
Keyword Scores
Deep Analysis
Innovations
- Proposes SGSimEval, a comprehensive multifaceted benchmark for automatic survey generation evaluation that integrates outline, content, and reference assessments.
- Combines LLM-based scoring with quantitative metrics to create a similarity-enhanced evaluation framework.
- Introduces human preference metrics that emphasize both inherent quality and similarity to human-generated surveys.
Methodology
SGSimEval evaluates automatic survey generation systems by assessing the outline, content, and references using a combination of LLM-based scoring and quantitative metrics, and incorporates human preference metrics that measure similarity to human surveys. Extensive experiments compare current ASG systems against human performance.
Key Results
Current ASG systems achieve human-comparable performance in outline generation but show significant room for improvement in content and reference generation; the proposed evaluation metrics maintain strong consistency with human assessments.