TaxoAlign: Scholarly Taxonomy Generation Using Language Models
TLDR
TaxoAlign generates scholarly taxonomies using a three-phase instruction-guided method, evaluated on a new benchmark of human-written survey taxonomies.
Reasoning
The paper introduces a novel benchmark (CS-TaxoBench) and a method (TaxoAlign) that outperforms baselines in structural and semantic alignment with human taxonomies. However, it is limited to taxonomy generation within computer science and does not address full survey generation or other scientific domains.
Read-first score
Read-first score 62.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 37.
Field roles
Rank sensitivity
Stability: volatile; rank range: 157.
Keyword Scores
Deep Analysis
Innovations
- CS-TaxoBench benchmark: 460 taxonomies from human-written survey papers plus 80 test taxonomies from conference surveys
- TaxoAlign: a three-phase topic-based instruction-guided method for scholarly taxonomy generation
- Stringent automated evaluation framework measuring structural alignment and semantic coherence against human-created taxonomies
Methodology
TaxoAlign is a three-phase topic-based instruction-guided method using language models. The CS-TaxoBench benchmark comprises 460 training taxonomies from human-written surveys and 80 test taxonomies from conference surveys. Evaluation uses automated metrics for structural alignment and semantic coherence, along with human evaluation studies, comparing against baselines.
Key Results
TaxoAlign consistently surpasses the baselines on nearly all metrics in both automated and human evaluations.