Large Language Models on Wikipedia-Style Survey Generation: an Evaluation in NLP Concepts
TLDR
Evaluates LLMs (GPT-4, etc.) for generating Wikipedia-style survey articles on NLP topics, finding GPT-4 best but with errors and evaluation bias.
Reasoning
Strengths: Clear focus on a specific task (survey generation) with automated benchmarks and human comparison. Weaknesses: Limited to NLP concepts, no real-world deployment or user study; GPT-4 still has factual errors and evaluation bias.
Read-first score
Read-first score 53.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 53.
Field roles
Rank sensitivity
Stability: volatile; rank range: 87.
Keyword Scores
Deep Analysis
Innovations
- First comprehensive evaluation of LLMs for generating Wikipedia-style survey articles in a specialized field (NLP)
- Curated benchmark of 99 NLP topics for survey generation
- Comparison of human and GPT-based evaluation revealing systematic bias in GPT evaluation
Methodology
Several LLMs (GPT-4, GPT-3.5, PaLM2, LLaMa2) were prompted to generate survey articles for 99 NLP topics. Automated benchmarks compared outputs to a ground truth, and both human and GPT-4 evaluations were conducted, followed by analysis of rating behavior and bias.
Key Results
GPT-4 outperformed other LLMs by 2–20% in automated metrics; GPT-generated surveys were more contemporary and accessible but occasionally missed details or contained factual errors; systematic bias was found in GPT-based evaluation.
Limitations
- GPT-4 occasionally lapses, missing details or introducing factual errors
- Systematic bias in using GPT-4 for evaluation compared to human ratings