Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

TaxoAlign: Scholarly Taxonomy Generation Using Language Models

arXiv 2025 62.8 method

TLDR

TaxoAlign generates scholarly taxonomies using a three-phase instruction-guided method, evaluated on a new benchmark of human-written survey taxonomies.

Reasoning

The paper introduces a novel benchmark (CS-TaxoBench) and a method (TaxoAlign) that outperforms baselines in structural and semantic alignment with human taxonomies. However, it is limited to taxonomy generation within computer science and does not address full survey generation or other scientific domains.

Read-first score

Read-first score 62.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 37.

Methodology quality 25%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,benchmark,evaluation,metric,result

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Reproducibility 25%
81

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=code,github

Topical relevance 42%
30.8

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

FrontierMethodology anchorReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 157.

Keyword Scores

survey generation
7
AI for scientific research
6
automated research
5
research automation
5
literature review agent
4
automated scientific discovery
3
AI scientist
2
paper writing agent
2
scientific discovery agent
2
autonomous research agent
1
automated experimentation
0
experiment design agent
0

Deep Analysis

Innovations

  • CS-TaxoBench benchmark: 460 taxonomies from human-written survey papers plus 80 test taxonomies from conference surveys
  • TaxoAlign: a three-phase topic-based instruction-guided method for scholarly taxonomy generation
  • Stringent automated evaluation framework measuring structural alignment and semantic coherence against human-created taxonomies

Methodology

TaxoAlign is a three-phase topic-based instruction-guided method using language models. The CS-TaxoBench benchmark comprises 460 training taxonomies from human-written surveys and 80 test taxonomies from conference surveys. Evaluation uses automated metrics for structural alignment and semantic coherence, along with human evaluation studies, comparing against baselines.

Key Results

TaxoAlign consistently surpasses the baselines on nearly all metrics in both automated and human evaluations.

Tags

CL