DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Generation
TLDR
DeepSurvey is an agentic system for automated survey generation that improves analytical depth and citation reliability using full-text analysis, cross-paper modeling, and evidence-constrained citation.
Reasoning
The paper presents a well-motivated approach addressing key limitations in automated survey generation, with strong empirical results including expert preference over human-written surveys. However, its focus is narrow (survey generation) and may not generalize to broader scientific discovery tasks without full-text access.
Read-first score
Read-first score 66.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 83.
Field roles
Rank sensitivity
Stability: volatile; rank range: 48.
Keyword Scores
Deep Analysis
Innovations
- Extracting structured keynotes from full-text papers to enhance analytical depth
- Modeling cross-paper relationships through clustering and comparative analysis
- Integrating code-repository analysis to recover implementation-level details
- Combining citation-graph expansion with hybrid filtering for topic-focussed retrieval
- Enforcing evidence-constrained citation assignment
- Deploying multi-granularity agentic refinement to validate citation-claim alignment
Methodology
DeepSurvey is an agentic system that enhances depth by extracting structured keynotes from full-text papers, modeling cross-paper relationships via clustering and comparative analysis, and integrating code-repository analysis. For citation reliability, it uses citation-graph expansion with hybrid filtering, evidence-constrained citation assignment, and multi-granularity agentic refinement to validate citation-claim alignment.
Key Results
DeepSurvey achieves the highest content score (8.644/10), citation quality gains of 12.3% recall and 9.3% precision over the strongest baseline, robust cross-domain generalization (0.14 vs 0.22 to 0.69 CS-to-non-CS drop), and is preferred over human-written surveys by domain experts (83.3% overall quality, 100% content depth).