Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle
TLDR
A benchmark series evaluating frontier LLMs and agentic systems on research lifecycle tasks, showing they still fall short of human researchers.
Reasoning
The paper introduces a novel benchmark (AARRI-Bench) that focuses on nuanced research behavior rather than macro execution, with extensive experiments across models. However, it only covers the first benchmark in the series and does not detail specific task types or limitations beyond the reported 68.3% success rate.
Read-first score
Read-first score 63.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.
Field roles
Rank sensitivity
Stability: volatile; rank range: 101.
Keyword Scores
Deep Analysis
Innovations
- Conceptualizes the AARR benchmark series that evaluates agents on granular research scenarios requiring professionalism, thoroughness, and nuanced reasoning, shifting focus from macro-level execution.
- Introduces AARRI-Bench, the first benchmark in the series, specifically targeting research intern-level tasks to assess whether agents can emulate human researcher qualities.
Methodology
The authors designed AARRI-Bench, a benchmark of granular research tasks that test agents' ability to exhibit professionalism, thoroughness, and nuanced reasoning. They evaluated multiple frontier models and agentic harnesses (e.g., Mini-SWE-Agent with Claude Opus 4.7) by measuring success rate on these tasks.
Key Results
The best configuration, Mini-SWE-Agent with Claude Opus 4.7, achieved only a 68.3% success rate, with agents frequently overlooking subtle yet critical details that are obvious to human researchers.
Limitations
- Agents exhibit significant limitations in field sensitivity, research ethics, and nuanced scientific judgment.
- Current frontier agents cannot fully replace human researchers.
- The benchmark is limited to research intern-level scenarios, not covering the full range of research activities.
- Developing researcher-like AI requires further exploration of research behavior beyond complex scaffolding.