Awesome Auto Research Hub 论文 · 数据集 · 项目
← 返回论文列表

Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents

arXiv 2026 47.7 method

TLDR

AutoSciRub is an evaluation-first framework that induces task-specific executable rubrics before research execution to guide autonomous agents, improving performance by 2-3 points on ResearchClawBench.

评分理由

The paper addresses the important problem of underspecified research tasks and proposes a novel rubric-induction method to make implicit requirements explicit. Its strengths include a clear framework and consistent empirical gains across multiple LLMs and harnesses. Limitations include reliance on a single benchmark and partial results on a 20-task subset.

Read-first 评分解释

综合优先阅读分 47.7,由主题、引用、图谱、方法、可复现性和近期性等信号加权得到。 原始总分保留为 66。

近期性 8%
100

使用温和的时间衰减,让近期论文更容易浮现,同时保留较早基础工作的价值。 年份:2026

方法质量 25%
60

检查可见的摘要与分析字段,寻找实验、数据集、基线、指标和局限性等方法证据。 命中信号:分析、评估、实验、结果

可复现性 25%
50

检查链接和可见文本中的论文、代码、数据集、工件与仓库信号。 论文:有;代码:无;数据:无;命中信号:工件、代码、GitHub

主题相关性 42%
28.6

将配置的研究关键词与标题、摘要、标签和分析文本进行匹配。 匹配关键词数:8

研究版图角色

前沿论文

排序敏感性

稳定性:volatile;排名波动范围:45。

关键词评分

autonomous research agent
8
AI for scientific research
8
automated research
7
research automation
7
scientific discovery agent
7
automated scientific discovery
6
AI scientist
5
automated experimentation
5
literature review agent
4
experiment design agent
4
paper writing agent
4
survey generation
1

深度分析

创新点

  • Proposes AutoSciRub, an evaluation-first framework that induces a task-specific executable rubric before research execution.
  • Decomposes underspecified research instructions into atomic scientific goals grounded in relevant literature and task-visible data.
  • Synthesizes specific, actionable, and verifiable criteria to make implicit experimental and evidential requirements explicit.
  • Uses rubric-guided criterion-level verification and iterative revision to identify unmet criteria and refine the research report and supporting artifacts.

方法

AutoSciRub first decomposes an underspecified instruction into atomic scientific goals, grounds them in literature and task-visible data, and synthesizes verifiable criteria into an executable rubric. The rubric then guides research execution, criterion-level verification, and iterative revision of the report and artifacts. The approach is evaluated on ResearchClawBench and a randomly sampled 20-task subset of AstaBench E2E Discovery across multiple backbone LLMs and agent harnesses, using average score gains and task completion as metrics.

关键结果

On ResearchClawBench, AutoSciRub improved all tested configurations, with average gains of 2.08 points across three backbone LLMs under a fixed Codex harness and 2.95 points across three agent harnesses using a fixed DeepSeek-V4-Flash backbone. On a 20-task AstaBench E2E Discovery subset, it achieved an average improvement of 16.8 points across three agent harnesses while maintaining or increasing the number of successfully completed tasks.

局限性

  • The AstaBench E2E Discovery evaluation is based on a randomly sampled 20-task subset rather than the full benchmark.

标签