Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

SciSage: A Multi-Agent Framework for High-Quality Scientific Survey Generation

arXiv 2025 60.9 method

TLDR

SciSage is a multi-agent framework with a hierarchical Reflector for generating high-quality scientific surveys, outperforming baselines on coherence and citation accuracy.

Reasoning

The paper introduces a novel multi-agent architecture with reflective evaluation at multiple levels, and releases a curated benchmark (SurveyScope). Strengths include clear methodology and quantitative gains; weaknesses are mixed human evaluation results and limited domain scope.

Read-first score

Read-first score 60.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 61.

Methodology quality 25%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,baseline,benchmark,evaluation,metric

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Topical relevance 42%
50.8

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 36.

Keyword Scores

survey generation
10
literature review agent
9
paper writing agent
8
AI for scientific research
7
autonomous research agent
6
research automation
6
automated research
5
automated scientific discovery
4
AI scientist
3
scientific discovery agent
3
automated experimentation
0
experiment design agent
0

Deep Analysis

Innovations

  • Multi-agent framework with a reflect-when-you-write paradigm
  • Hierarchical Reflector agent that evaluates drafts at outline, section, and document levels
  • Specialized agents for query interpretation, content retrieval, and refinement
  • SurveyScope benchmark: 46 high-impact papers (2020-2025) across 11 computer science domains with strict recency and citation-based quality controls

Methodology

SciSage is a multi-agent system where a hierarchical Reflector agent critiques drafts at multiple granularities, collaborating with agents for query interpretation, retrieval, and refinement. It is evaluated on the SurveyScope benchmark against LLM x MapReduce-V2 and AutoSurvey baselines using automatic metrics (document coherence, citation F1) and human evaluation.

Key Results

SciSage outperforms baselines by +1.73 points in document coherence and +32% in citation F1. Human evaluation shows 3 wins vs. 7 losses against human-written surveys, with strengths in topical breadth and retrieval efficiency.

Limitations

  • Human evaluation reveals 7 losses vs. 3 wins against human-written surveys, indicating remaining quality gaps.
  • SurveyScope is limited to 46 papers in computer science, potentially restricting generalizability to other domains.

Tags

AIIR