Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

arXiv '24 2023 34.2 benchmark

Read-first score

Read-first score 34.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals.

Reproducibility 33%
73

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=github

Recency 11%
65.1

Uses a gentle age decay so recent papers surface without erasing older foundations. 2023

Topical relevance 56%
4.8

Matches configured research keywords against title, abstract, tags, and analysis text. matched=1

Field roles

Reproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 105.

Tags