Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

Large Language Models on Wikipedia-Style Survey Generation: an Evaluation in NLP Concepts

arXiv 2023 53.8 method

TLDR

Evaluates LLMs (GPT-4, etc.) for generating Wikipedia-style survey articles on NLP topics, finding GPT-4 best but with errors and evaluation bias.

Reasoning

Strengths: Clear focus on a specific task (survey generation) with automated benchmarks and human comparison. Weaknesses: Limited to NLP concepts, no real-world deployment or user study; GPT-4 still has factual errors and evaluation bias.

Read-first score

Read-first score 53.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 53.

Methodology quality 25%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,benchmark,evaluation,metric

Recency 8%
65.1

Uses a gentle age decay so recent papers surface without erasing older foundations. 2023

Topical relevance 42%
44.2

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

Methodology anchor

Rank sensitivity

Stability: volatile; rank range: 87.

Keyword Scores

survey generation
10
literature review agent
9
paper writing agent
8
AI for scientific research
7
automated research
6
research automation
6
autonomous research agent
3
AI scientist
2
automated scientific discovery
1
scientific discovery agent
1
automated experimentation
0
experiment design agent
0

Deep Analysis

Innovations

  • First comprehensive evaluation of LLMs for generating Wikipedia-style survey articles in a specialized field (NLP)
  • Curated benchmark of 99 NLP topics for survey generation
  • Comparison of human and GPT-based evaluation revealing systematic bias in GPT evaluation

Methodology

Several LLMs (GPT-4, GPT-3.5, PaLM2, LLaMa2) were prompted to generate survey articles for 99 NLP topics. Automated benchmarks compared outputs to a ground truth, and both human and GPT-4 evaluations were conducted, followed by analysis of rating behavior and bias.

Key Results

GPT-4 outperformed other LLMs by 2–20% in automated metrics; GPT-generated surveys were more contemporary and accessible but occasionally missed details or contained factual errors; systematic bias was found in GPT-based evaluation.

Limitations

  • GPT-4 occasionally lapses, missing details or introducing factual errors
  • Systematic bias in using GPT-4 for evaluation compared to human ratings

Tags

CL