MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?
TLDR
Introduces MLRC-Bench, a benchmark evaluating language agents on machine learning research challenges with objective metrics, showing significant gaps.
Reasoning
Strengths include rigorous objective evaluation and identification of misalignment with LLM-judged innovation. Weaknesses are limited task scope (7 tasks) and lack of detailed agent comparisons beyond one best performer.
Read-first score
Read-first score 53.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 42.
Field roles
Rank sensitivity
Stability: volatile; rank range: 65.
Keyword Scores
Deep Analysis
Innovations
- Introduces MLRC-Bench, a benchmark that evaluates language agents on proposing and implementing novel ML research methods using objective metrics from competition tasks, unlike prior LLM-as-a-judge approaches.
- Reveals a misalignment between LLM-judged innovation and actual task performance in ML research challenges.
- Proposes a dynamic benchmark that can incorporate new ML competitions for ongoing rigorous evaluation.
Methodology
MLRC-Bench consists of 7 curated ML research competition tasks requiring novel methodologies. Language agents are evaluated on their ability to propose and implement research methods, with performance measured by objective competition metrics and compared against baselines and top human participant scores.
Key Results
The best agent (gemini-exp-1206 under MLAB) closed only 9.3% of the gap between baseline and top human scores, and LLM-judged innovation was misaligned with actual performance.