
数据集大全
数据集
按表示方法和任务分组的数据集。




AlphaResearch
LLMs have made significant progress in complex but easy-to-verify problems, yet they still struggle with discovering the unknown. In this paper, we present \textbf{AlphaResearch}, an autonomous research agent designed to discover new algori...






























DeepScholar-Bench
Derived from paper: DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis

Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
Derived from paper: Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows







ReportBench
Derived from paper: ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks

SciReplicate-Bench
Derived from paper: SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers



DiscoveryBench
Derived from paper: DiscoveryBench: Towards Data-Driven Discovery with Large Language Models


MLAgentBench
Derived from paper: MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation


AutoResearchBench
Derived from paper: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery

Process-Oriented Evaluation of AI-Assisted Scientific Writing
Derived from paper: Process-Oriented Evaluation of AI-Assisted Scientific Writing

ResearchClawBench
Derived from paper: ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

SciFlow-Bench
Derived from paper: SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing

SciNetBench
Derived from paper: SciNetBench: A Relation-Aware Benchmark for Scientific Literature Retrieval Agents


ResearchBench
Derived from paper: ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition



LiveIdeaBench
Derived from paper: LiveIdeaBench: Evaluating LLMs' Scientific Creativity and Idea Generation with Minimal Context



PseudoBench
Derived from paper: PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience

ScholarGym
Derived from paper: ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research


Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper
Derived from paper: Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper


Can LLMs Generate Novel Research Ideas? A Large Scale Human Study with 100+ NLP Researchers
Derived from paper: Can LLMs Generate Novel Research Ideas? A Large Scale Human Study with 100+ NLP Researchers

alexshengzhili/ai-scientist-blog-assets
Hugging Face dataset: alexshengzhili/ai-scientist-blog-assets
jablonkagroup/rise_ai_scientists
Corral – Rise of AI Scientists Bibliometric evidence for the rise of AI scientists in chemistry and materials science relative to general AI for chemistry 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the bibliometric evidence for the growing relevance and impact of AI scientists in chemistry and materials science compared to general AI… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/rise_ai_scientists.
automated-research-group/llama2_7b-arc_hard
Dataset Card for "llama2_7b-arc_hard" More Information needed
automated-research-group/llama2_7b-arc_hard-results_playing
Dataset Card for "llama2_7b-arc_hard-results_playing" More Information needed
automated-research-group/boolq
Dataset Card for "boolq" More Information needed
automated-research-group/gpt2-winogrande_base
Dataset Card for "gpt2-winogrande_base" More Information needed
automated-research-group/gpt2-winogrande_inverted_option
Dataset Card for "gpt2-winogrande_inverted_option" More Information needed
automated-research-group/llama2_7b_chat-agieval-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-agieval-results
automated-research-group/llama2_7b_chat-arc_challenge-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-arc_challenge-results
automated-research-group/llama2_7b_chat-arc_easy-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-arc_easy-results
automated-research-group/llama2_7b_chat-boolq
Dataset Card for "llama2_7b_chat-boolq" More Information needed
automated-research-group/llama2_7b_chat-boolq-results
Dataset Card for "llama2_7b_chat-boolq-results" More Information needed
automated-research-group/llama2_7b_chat-boolq-results_jacksee
Dataset Card for "llama2_7b_chat-boolq-results_jacksee" More Information needed
automated-research-group/llama2_7b_chat-commonsense_qa-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-commonsense_qa-results
automated-research-group/llama2_7b_chat-hellaswag_0_label
Hugging Face dataset: automated-research-group/llama2_7b_chat-hellaswag_0_label
automated-research-group/llama2_7b_chat-hellaswag_0_label-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-hellaswag_0_label-results
automated-research-group/llama2_7b_chat-hellaswag-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-hellaswag-results
automated-research-group/llama2_7b_chat-openbookqa-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-openbookqa-results
automated-research-group/llama2_7b_chat-piqa-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-piqa-results
automated-research-group/llama2_7b_chat-siqa
Hugging Face dataset: automated-research-group/llama2_7b_chat-siqa
automated-research-group/llama2_7b_chat-siqa-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-siqa-results
automated-research-group/llama2_7b_chat-winogrande-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-winogrande-results
automated-research-group/phi-boolq-results
Dataset Card for "phi-boolq-results" More Information needed
automated-research-group/phi-boolq-results_playing
Dataset Card for "phi-boolq-results_playing" More Information needed
automated-research-group/phi-winogrande_base
Dataset Card for "phi-winogrande_base" More Information needed
automated-research-group/phi-winogrande_inverted_option
Dataset Card for "phi-winogrande_inverted_option" More Information needed
automated-research-group/phi-winogrande_inverted_option-results
Dataset Card for "phi-winogrande_inverted_option-results" More Information needed
automated-research-group/winogrande
Hugging Face dataset: automated-research-group/winogrande
automated-research-group/winogrande_inverted_option
Dataset Card for "winogrande_inverted_option" More Information needed