MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility
TLDR
Introduces MLReplicate, a benchmark evaluating autonomous research systems on ML reproducibility using ICML 2025 papers, finding automated reviews unreliable and cost unrelated to quality.
Reasoning
Strengths include a novel dual-protocol evaluation combining automated and human review, revealing critical flaws in current systems. Weaknesses are the narrow focus on ML reproducibility and limited sample size of 6 systems and 45 manuscripts.
Read-first score
Read-first score 64.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 75.
Field roles
Rank sensitivity
Stability: volatile; rank range: 32.
Keyword Scores
Deep Analysis
Innovations
- MLReplicate benchmark: an end-to-end evaluation framework for autonomous research systems on machine learning reproducibility, built from ICML 2025 outstanding papers reformulated into standardized input specifications.
- Dual-protocol evaluation combining automated conference-style review and structured expert human evaluation, while tracking computational cost, runtime, and human intervention.
- Demonstration that autonomous research workflow design matters more than compute scale, as the cheapest system outperforms the most resource-intensive system in human evaluation despite a 38-fold difference in input tokens.
Methodology
MLReplicate was constructed from ICML 2025 outstanding papers reformulated into standardized input specifications. Six state-of-the-art autonomous research systems (AI SCIENTIST-V1, AI SCIENTIST-V2, AGENT LABORATORY, CYCLERESEARCHER, AI RESEARCHER, TINY SCIENTIST) generated 45 manuscripts (3 failed). Outputs were assessed via automated conference-style review and structured expert human evaluation, with metrics including computational cost, runtime, and human intervention.
Key Results
Automated review accepted 10 of 37 valid submissions, but human reviewers identified methodological flaws, hallucinated results, and reproducibility failures in all systems; 59% of accepted automated reviews contained fabricated or unsupported claims. The cheapest system outperformed the most resource-intensive system in human evaluation, showing workflow design matters more than compute scale.