Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?

arXiv 2025 53.8 benchmark

TLDR

Introduces MLRC-Bench, a benchmark evaluating language agents on machine learning research challenges with objective metrics, showing significant gaps.

Reasoning

Strengths include rigorous objective evaluation and identification of misalignment with LLM-judged innovation. Weaknesses are limited task scope (7 tasks) and lack of detailed agent comparisons beyond one best performer.

Read-first score

Read-first score 53.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 42.

Methodology quality 25%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,baseline,benchmark,evaluation,metric

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code

Topical relevance 42%
35

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 65.

Keyword Scores

autonomous research agent
6
research automation
6
automated research
5
AI for scientific research
5
scientific discovery agent
5
automated scientific discovery
4
AI scientist
3
automated experimentation
3
experiment design agent
2
literature review agent
1
survey generation
1
paper writing agent
1

Deep Analysis

Innovations

  • Introduces MLRC-Bench, a benchmark that evaluates language agents on proposing and implementing novel ML research methods using objective metrics from competition tasks, unlike prior LLM-as-a-judge approaches.
  • Reveals a misalignment between LLM-judged innovation and actual task performance in ML research challenges.
  • Proposes a dynamic benchmark that can incorporate new ML competitions for ongoing rigorous evaluation.

Methodology

MLRC-Bench consists of 7 curated ML research competition tasks requiring novel methodologies. Language agents are evaluated on their ability to propose and implement research methods, with performance measured by objective competition metrics and compared against baselines and top human participant scores.

Key Results

The best agent (gemini-exp-1206 under MLAB) closed only 9.3% of the gap between baseline and top human scores, and LLM-judged innovation was misaligned with actual performance.

Tags

AI