Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle

arXiv 2026 63.2 method

TLDR

A benchmark series evaluating frontier LLMs and agentic systems on research lifecycle tasks, showing they still fall short of human researchers.

Reasoning

The paper introduces a novel benchmark (AARRI-Bench) that focuses on nuanced research behavior rather than macro execution, with extensive experiments across models. However, it only covers the first benchmark in the series and does not detail specific task types or limitations beyond the reported 68.3% success rate.

Read-first score

Read-first score 63.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 48.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,experiment,result

Reproducibility 25%
73

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=github

Topical relevance 42%
40

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

FrontierMethodology anchorReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 101.

Keyword Scores

autonomous research agent
8
AI for scientific research
7
research automation
7
automated research
6
automated scientific discovery
4
scientific discovery agent
4
AI scientist
3
automated experimentation
3
literature review agent
2
experiment design agent
2
survey generation
1
paper writing agent
1

Deep Analysis

Innovations

  • Conceptualizes the AARR benchmark series that evaluates agents on granular research scenarios requiring professionalism, thoroughness, and nuanced reasoning, shifting focus from macro-level execution.
  • Introduces AARRI-Bench, the first benchmark in the series, specifically targeting research intern-level tasks to assess whether agents can emulate human researcher qualities.

Methodology

The authors designed AARRI-Bench, a benchmark of granular research tasks that test agents' ability to exhibit professionalism, thoroughness, and nuanced reasoning. They evaluated multiple frontier models and agentic harnesses (e.g., Mini-SWE-Agent with Claude Opus 4.7) by measuring success rate on these tasks.

Key Results

The best configuration, Mini-SWE-Agent with Claude Opus 4.7, achieved only a 68.3% success rate, with agents frequently overlooking subtle yet critical details that are obvious to human researchers.

Limitations

  • Agents exhibit significant limitations in field sensitivity, research ethics, and nuanced scientific judgment.
  • Current frontier agents cannot fully replace human researchers.
  • The benchmark is limited to research intern-level scenarios, not covering the full range of research activities.
  • Developing researcher-like AI requires further exploration of research behavior beyond complex scaffolding.

Tags

AI