Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

PaperBench: Evaluating AI's Ability to Replicate AI Research

arXiv '25 2025 34 benchmark

Read-first score

Read-first score 34, weighted from topical fit, citation, graph, method, reproducibility, and recency signals.

Recency 11%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Reproducibility 33%
73

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=github

Topical relevance 56%
0

Matches configured research keywords against title, abstract, tags, and analysis text. matched=0

Field roles

FrontierReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 94.

Tags