Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility

arXiv 2026 64.4 method

TLDR

Introduces MLReplicate, a benchmark evaluating autonomous research systems on ML reproducibility using ICML 2025 papers, finding automated reviews unreliable and cost unrelated to quality.

Reasoning

Strengths include a novel dual-protocol evaluation combining automated and human review, revealing critical flaws in current systems. Weaknesses are the narrow focus on ML reproducibility and limited sample size of 6 systems and 45 manuscripts.

Read-first score

Read-first score 64.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 75.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,experiment,metric,result

Topical relevance 42%
62.5

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 32.

Keyword Scores

AI scientist
9
autonomous research agent
9
automated research
8
AI for scientific research
8
paper writing agent
8
research automation
8
automated scientific discovery
7
scientific discovery agent
7
automated experimentation
5
experiment design agent
3
literature review agent
2
survey generation
1

Deep Analysis

Innovations

  • MLReplicate benchmark: an end-to-end evaluation framework for autonomous research systems on machine learning reproducibility, built from ICML 2025 outstanding papers reformulated into standardized input specifications.
  • Dual-protocol evaluation combining automated conference-style review and structured expert human evaluation, while tracking computational cost, runtime, and human intervention.
  • Demonstration that autonomous research workflow design matters more than compute scale, as the cheapest system outperforms the most resource-intensive system in human evaluation despite a 38-fold difference in input tokens.

Methodology

MLReplicate was constructed from ICML 2025 outstanding papers reformulated into standardized input specifications. Six state-of-the-art autonomous research systems (AI SCIENTIST-V1, AI SCIENTIST-V2, AGENT LABORATORY, CYCLERESEARCHER, AI RESEARCHER, TINY SCIENTIST) generated 45 manuscripts (3 failed). Outputs were assessed via automated conference-style review and structured expert human evaluation, with metrics including computational cost, runtime, and human intervention.

Key Results

Automated review accepted 10 of 37 valid submissions, but human reviewers identified methodological flaws, hallucinated results, and reproducibility failures in all systems; 59% of accepted automated reviews contained fabricated or unsupported claims. The cheapest system outperformed the most resource-intensive system in human evaluation, showing workflow design matters more than compute scale.

Tags

LG