
数据集大全
数据集
按表示方法和任务分组的数据集。




AlphaResearch
LLMs have made significant progress in complex but easy-to-verify problems, yet they still struggle with discovering the unknown. In this paper, we present \textbf{AlphaResearch}, an autonomous research agent designed to discover new algori...
































DeepScholar-Bench
Derived from paper: DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis

Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
Derived from paper: Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows





社会科学研究设计追踪与评估
对社会科学中因果研究设计的可靠评估对于循证政策制定至关重要,但迄今为止完全依赖人工专家分析。我们提出了自动研究设计追踪与评估(ARDTrA),该任务涉及检测论文中使用的研究设计并评估其应用质量。我们创建了一个由专家标注的论文数据集,涵盖六类反事实研究设计,并使用基于RAG的多轮对话流水线对该任务进行评估。在四种检索策略、四种LLM和六种嵌入模型的实验下,我们发现段落长度是性能的主要驱动因素,解释了52%-66%的方差。按研究设计进行的分析还表明,人类与机器的难度并不一致:系统最难处理的设计并非专家标注者分歧最大的设计,这表明任务难度存在两个独立的来源。



ReportBench
Derived from paper: ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks

SciReplicate-Bench
Derived from paper: SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers




DiscoveryBench
Derived from paper: DiscoveryBench: Towards Data-Driven Discovery with Large Language Models


ASI-Bench
超级人工智能(ASI)要求人工智能超越对现有知识的掌握,转向探索未知、创造新知识,并将新想法转化为可验证的成果。然而,当今人工智能系统的能力在很大程度上仍建立在学习、压缩和应用人类现有知识的基础之上。因此,现有基准主要测试人工智能能否基于已学知识给出正确答案,或能否在大量人类指导下完成任务。为此,我们推出 ASI-Bench,这是首个在通用研究领域内联合评估人工智能系统创新探索能力与自主科学执行能力的基准,也是首个在同一研究项目中逐步撤除人类方法论指导、以测试人工智能在多大程度上能够独立推进的基准。ASI-Bench 由 40 余位专家耗费 31,000 多个人工小时构建,包含涵盖 11 个科学领域的 60 个项目级研究任务,并逐步减少方法论指导,以测试人工智能能否独立选择方法、开展研究并产出可验证的结果。所有任务均经过专家评审、AI 辅助审计、沙盒执行和评分者验证。在 18 种最先进的智能体-模型配置中,平均得分从提供完整方法论指导时的 50.91 分,降至仅指定方法时的 29.10 分,再降至智能体必须自行确定方法时的 26.62 分。这一急剧下降表明,当前系统仍严重依赖人类指导,距离自主开展端到端的项目级科学研究还有很大差距。ASI-Bench 向全球开放。我们邀请各地的研究者和开发者贡献新任务、挑战当今 AI 的极限,并共同加速人类迈向超级人工智能的集体进程,网址为 https://asibench.apexin.ai/submit。

MLAgentBench
Derived from paper: MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation



AutoResearchBench
Derived from paper: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery

Process-Oriented Evaluation of AI-Assisted Scientific Writing
Derived from paper: Process-Oriented Evaluation of AI-Assisted Scientific Writing

ResearchClawBench
Derived from paper: ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

SciFlow-Bench
Derived from paper: SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing

SciNetBench
Derived from paper: SciNetBench: A Relation-Aware Benchmark for Scientific Literature Retrieval Agents


ResearchBench
Derived from paper: ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition



LiveIdeaBench
Derived from paper: LiveIdeaBench: Evaluating LLMs' Scientific Creativity and Idea Generation with Minimal Context



PseudoBench
Derived from paper: PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience

ScholarGym
Derived from paper: ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research


Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper
Derived from paper: Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper


Can LLMs Generate Novel Research Ideas? A Large Scale Human Study with 100+ NLP Researchers
Derived from paper: Can LLMs Generate Novel Research Ideas? A Large Scale Human Study with 100+ NLP Researchers

alexshengzhili/ai-scientist-blog-assets
Hugging Face dataset: alexshengzhili/ai-scientist-blog-assets
jablonkagroup/rise_ai_scientists
Corral – Rise of AI Scientists Bibliometric evidence for the rise of AI scientists in chemistry and materials science relative to general AI for chemistry 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the bibliometric evidence for the growing relevance and impact of AI scientists in chemistry and materials science compared to general AI… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/rise_ai_scientists.
automated-research-group/llama2_7b-arc_hard
Dataset Card for "llama2_7b-arc_hard" More Information needed
automated-research-group/llama2_7b-arc_hard-results_playing
Dataset Card for "llama2_7b-arc_hard-results_playing" More Information needed
automated-research-group/boolq
Dataset Card for "boolq" More Information needed
automated-research-group/gpt2-winogrande_base
Dataset Card for "gpt2-winogrande_base" More Information needed
automated-research-group/gpt2-winogrande_inverted_option
Dataset Card for "gpt2-winogrande_inverted_option" More Information needed
automated-research-group/llama2_7b_chat-agieval-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-agieval-results
automated-research-group/llama2_7b_chat-arc_challenge-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-arc_challenge-results
automated-research-group/llama2_7b_chat-arc_easy-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-arc_easy-results
automated-research-group/llama2_7b_chat-boolq
Dataset Card for "llama2_7b_chat-boolq" More Information needed
automated-research-group/llama2_7b_chat-boolq-results
Dataset Card for "llama2_7b_chat-boolq-results" More Information needed
automated-research-group/llama2_7b_chat-boolq-results_jacksee
Dataset Card for "llama2_7b_chat-boolq-results_jacksee" More Information needed
automated-research-group/llama2_7b_chat-commonsense_qa-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-commonsense_qa-results
automated-research-group/llama2_7b_chat-hellaswag_0_label
Hugging Face dataset: automated-research-group/llama2_7b_chat-hellaswag_0_label
automated-research-group/llama2_7b_chat-hellaswag_0_label-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-hellaswag_0_label-results
automated-research-group/llama2_7b_chat-hellaswag-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-hellaswag-results
automated-research-group/llama2_7b_chat-openbookqa-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-openbookqa-results
automated-research-group/llama2_7b_chat-piqa-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-piqa-results
automated-research-group/llama2_7b_chat-siqa
Hugging Face dataset: automated-research-group/llama2_7b_chat-siqa
automated-research-group/llama2_7b_chat-siqa-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-siqa-results
automated-research-group/llama2_7b_chat-winogrande-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-winogrande-results
automated-research-group/phi-boolq-results
Dataset Card for "phi-boolq-results" More Information needed
automated-research-group/phi-boolq-results_playing
Dataset Card for "phi-boolq-results_playing" More Information needed
automated-research-group/phi-winogrande_base
Dataset Card for "phi-winogrande_base" More Information needed
automated-research-group/phi-winogrande_inverted_option
Dataset Card for "phi-winogrande_inverted_option" More Information needed
automated-research-group/phi-winogrande_inverted_option-results
Dataset Card for "phi-winogrande_inverted_option-results" More Information needed
automated-research-group/winogrande
Hugging Face dataset: automated-research-group/winogrande
automated-research-group/winogrande_inverted_option
Dataset Card for "winogrande_inverted_option" More Information needed