Awesome Auto Research Hub 论文 · 数据集 · 项目
← 返回论文列表

迈向可审计的AI科学家:面向LLM智能体的假设演化协议

arXiv 2026 54.1 method

TLDR

提出假设演化协议(HEP),使LLM智能体在科学发现中的假设生成、评估和演化过程可审计。

评分理由

The paper addresses a clear gap in auditability of LLM-based scientific agents and provides a structured protocol. However, the evaluation is limited to materials-science tasks without explicit mention of real-world datasets or benchmarks, and the abstract lacks details on comparative baselines or limitations.

Read-first 评分解释

综合优先阅读分 54.1,由主题、引用、图谱、方法、可复现性和近期性等信号加权得到。 原始总分保留为 67。

近期性 8%
100

使用温和的时间衰减,让近期论文更容易浮现,同时保留较早基础工作的价值。 年份:2026

方法质量 25%
60

检查可见的摘要与分析字段,寻找实验、数据集、基线、指标和局限性等方法证据。 命中信号:评估、结果

主题相关性 42%
55.8

使用现有 LLM 关键词相关性评分,并归一化到 0-100。 关键词:AI scientist、automated scientific discovery、autonomous research agent、automated research、literature review agent、survey generation、automated experimentation、experiment design agent、AI for scientific research、paper writing agent、research automation、scientific discovery agent

可复现性 25%
30

检查链接和可见文本中的论文、代码、数据集、工件与仓库信号。 论文:有;代码:无;数据:无;命中信号:无

研究版图角色

前沿论文

排序敏感性

稳定性:volatile;排名波动范围:68。

关键词评分

AI scientist
9
AI for scientific research
9
scientific discovery agent
9
automated scientific discovery
8
autonomous research agent
8
automated research
7
research automation
7
automated experimentation
5
experiment design agent
4
literature review agent
1
survey generation
0
paper writing agent
0

深度分析

创新点

  • 假设演化协议(HEP),将LLM智能体的假设生成、评估和演化变为显式、可审计的操作

方法

HEP是一种框架,将LLM智能体的科学推理结构化为假设生成、评估和演化的显式步骤。该方法在材料科学研究任务上进行评估,将配备HEP的智能体与规划型智能体进行比较。

关键结果

配备HEP的智能体能够执行假设-测试-证据-信念循环,跨研究问题泛化,并且随着基础LLM能力的增强,更充分地利用该协议。

标签