
Dataset zoo
Datasets
Datasets grouped by representation and task.


SciIntegrity-Bench
AI scientist systems are increasingly deployed for autonomous research, yet their academic integrity has never been systematically evaluated. We introduce SCIINTEGRITY-BENCH, the first benchmark designed around a dilemmatic evaluation parad...


AlphaResearch
LLMs have made significant progress in complex but easy-to-verify problems, yet they still struggle with discovering the unknown. In this paper, we present \textbf{AlphaResearch}, an autonomous research agent designed to discover new algori...

CycleResearcher
The automation of scientific discovery has been a long-standing goal within the research community, driven by the potential to accelerate knowledge creation. While significant progress has been made using commercial large language models (L...

BioKGBench
Pursuing artificial intelligence for biomedical science, a.k.a. AI Scientist, draws increasing attention, where one common approach is to build a copilot agent driven by Large Language Models (LLMs). However, to evaluate such systems, peopl...


Dynamic Knowledge Exchange and Dual-diversity Review
Scientific progress increasingly relies on effective collaboration among researchers, a dynamic that large language models (LLMs) have only begun to emulate. While recent LLM-based scientist agents show promise in autonomous scientific disc...

When AI Does Science
Agentic AI "scientists" now use language models to search the literature, run analyses, and generate hypotheses. We evaluate KOSMOS, an autonomous AI scientist, on three problems in radiation biology using simple random-gene null benchmarks...

Act As a Real Researcher
As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experiment execution. Despite their evolution f...

DeepSurvey-Bench
The rapid development of automated scientific survey generation technology has made it increasingly important to establish a comprehensive benchmark to evaluate the quality of generated surveys.Nearly all existing evaluation benchmarks rely...




MegaScience
Scientific reasoning is critical for developing AI scientists and supporting human researchers in advancing the frontiers of natural science discovery. However, the open-source community has primarily focused on mathematics and coding while...

PaperSearchQA
Search agents are language models (LMs) that reason and search knowledge bases (or the web) to answer questions; recent methods supervise only the final answer accuracy using reinforcement learning with verifiable rewards (RLVR). Most RLVR ...

SurveyLens
Automatic Survey Generation (ASG) aims to produce comprehensive literature surveys by retrieving, organizing, and synthesizing academic papers. Despite rapid progress in specialized ASG frameworks and Deep Research agents, existing evaluati...

MLRC-Bench
We introduce MLRC-Bench, a benchmark designed to quantify how effectively language agents can tackle challenging Machine Learning (ML) Research Competitions, with a focus on open research problems that demand novel methodologies. Unlike pri...


Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
SDABench evaluates LLMs on six scientific data analysis capabilities across five domains, finding strengths in descriptive analysis but weaknesses in assumption selection and mechanistic reasoning.

AstroVisBench
Large Language Models (LLMs) are being explored for applications in scientific research, including their capabilities to synthesize literature, answer research questions, generate research ideas, and even conduct computational experiments. ...

When AI Co-Scientists Fail
Recent advances in large language models (LLMs) have fueled the vision of automated scientific discovery, often called AI Co-Scientists. To date, prior work casts these systems as generative co-authors responsible for crafting hypotheses, s...

Exploring the Limitations of kNN Noisy Feature Detection and Recovery for Self-Driving Labs
Self-driving laboratories (SDLs) have shown promise to accelerate materials discovery by integrating machine learning with automated experimental platforms. However, errors in the capture of input parameters may corrupt the features used to...



The Automated LLM Speedrunning Benchmark
Rapid advancements in large language models (LLMs) have the potential to assist in scientific progress. A critical capability toward this endeavor is the ability to reproduce existing work. To evaluate the ability of AI agents to reproduce ...



PaperBanana
Despite rapid advances in autonomous AI scientists powered by language models, generating publication-ready illustrations remains a labor-intensive bottleneck in the research workflow. To lift this burden, we introduce PaperBanana, an agent...

Joint discovery of governing partial differential equations from multi-source datasets by competitive optimization
Discovering governing equations directly from observational data is a key step towards interpretable scientific machine learning. Current data-driven approaches typically operate on a single dataset, inherently limiting their performance wh...

Understanding Usage and Engagement in AI-Powered Scientific Research Tools
AI-powered scientific research tools are rapidly being integrated into research workflows, yet the field lacks a clear lens into how researchers use these systems in real-world settings. We present and analyze the Asta Interaction Dataset, ...

Evolution Fine-Tuning
Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture? Large Language Models (LLMs) integrated into evolutionary search have recently produced state-of-the-art solutions on optimiza...

Unlocking the Visual Record of Materials Science
The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains locked away and inaccessible to AI at scale. The core difficulty is structural: most scientific figures are compound, with...

DeepScholar-Bench
Derived from paper: DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis

Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
Derived from paper: Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows



Beyond Drug Discovery
Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets. This combination yields strong benchmark metrics but limits transferability to domains structu...

An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics
The rigorous evaluation of the novelty of a scientific paper is, even for human scientists, a challenging task. With the increasing interest in AI scientists and AI involvement in scientific idea generation and paper writing, it also become...


AI Idea Bench 2025
Derived from paper: AI Idea Bench 2025: AI Research Idea Generation Benchmark

ReportBench
Derived from paper: ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks

SciReplicate-Bench
Derived from paper: SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers


PostTrainBench
Derived from paper: PostTrainBench: Can LLM Agents Automate LLM Post-Training?

DiscoveryBench
Derived from paper: DiscoveryBench: Towards Data-Driven Discovery with Large Language Models


MLAgentBench
Derived from paper: MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation


AutoResearchBench
Derived from paper: AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery

Process-Oriented Evaluation of AI-Assisted Scientific Writing
Derived from paper: Process-Oriented Evaluation of AI-Assisted Scientific Writing

ResearchClawBench
Derived from paper: ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

SciFlow-Bench
Derived from paper: SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing

SciNetBench
Derived from paper: SciNetBench: A Relation-Aware Benchmark for Scientific Literature Retrieval Agents


ResearchBench
Derived from paper: ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition
ClaimCheck
Derived from paper: ClaimCheck: How Grounded are LLM Critiques of Scientific Papers?

FrontierScience
Derived from paper: FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks


LiveIdeaBench
Derived from paper: LiveIdeaBench: Evaluating LLMs' Scientific Creativity and Idea Generation with Minimal Context



PseudoBench
Derived from paper: PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience

ScholarGym
Derived from paper: ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research


Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper
Derived from paper: Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper


Can LLMs Generate Novel Research Ideas? A Large Scale Human Study with 100+ NLP Researchers
Derived from paper: Can LLMs Generate Novel Research Ideas? A Large Scale Human Study with 100+ NLP Researchers

alexshengzhili/ai-scientist-blog-assets
Hugging Face dataset: alexshengzhili/ai-scientist-blog-assets
jablonkagroup/rise_ai_scientists
Corral – Rise of AI Scientists Bibliometric evidence for the rise of AI scientists in chemistry and materials science relative to general AI for chemistry 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the bibliometric evidence for the growing relevance and impact of AI scientists in chemistry and materials science compared to general AI… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/rise_ai_scientists.
automated-research-group/llama2_7b-arc_hard
Dataset Card for "llama2_7b-arc_hard" More Information needed
automated-research-group/llama2_7b-arc_hard-results_playing
Dataset Card for "llama2_7b-arc_hard-results_playing" More Information needed
automated-research-group/boolq
Dataset Card for "boolq" More Information needed
automated-research-group/gpt2-winogrande_base
Dataset Card for "gpt2-winogrande_base" More Information needed
automated-research-group/gpt2-winogrande_inverted_option
Dataset Card for "gpt2-winogrande_inverted_option" More Information needed
automated-research-group/llama2_7b_chat-agieval-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-agieval-results
automated-research-group/llama2_7b_chat-arc_challenge-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-arc_challenge-results
automated-research-group/llama2_7b_chat-arc_easy-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-arc_easy-results
automated-research-group/llama2_7b_chat-boolq
Dataset Card for "llama2_7b_chat-boolq" More Information needed
automated-research-group/llama2_7b_chat-boolq-results
Dataset Card for "llama2_7b_chat-boolq-results" More Information needed
automated-research-group/llama2_7b_chat-boolq-results_jacksee
Dataset Card for "llama2_7b_chat-boolq-results_jacksee" More Information needed
automated-research-group/llama2_7b_chat-commonsense_qa-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-commonsense_qa-results
automated-research-group/llama2_7b_chat-hellaswag_0_label
Hugging Face dataset: automated-research-group/llama2_7b_chat-hellaswag_0_label
automated-research-group/llama2_7b_chat-hellaswag_0_label-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-hellaswag_0_label-results
automated-research-group/llama2_7b_chat-hellaswag-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-hellaswag-results
automated-research-group/llama2_7b_chat-openbookqa-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-openbookqa-results
automated-research-group/llama2_7b_chat-piqa-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-piqa-results
automated-research-group/llama2_7b_chat-siqa
Hugging Face dataset: automated-research-group/llama2_7b_chat-siqa
automated-research-group/llama2_7b_chat-siqa-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-siqa-results
automated-research-group/llama2_7b_chat-winogrande-results
Hugging Face dataset: automated-research-group/llama2_7b_chat-winogrande-results
automated-research-group/phi-boolq-results
Dataset Card for "phi-boolq-results" More Information needed
automated-research-group/phi-boolq-results_playing
Dataset Card for "phi-boolq-results_playing" More Information needed
automated-research-group/phi-winogrande_base
Dataset Card for "phi-winogrande_base" More Information needed
automated-research-group/phi-winogrande_inverted_option
Dataset Card for "phi-winogrande_inverted_option" More Information needed
automated-research-group/phi-winogrande_inverted_option-results
Dataset Card for "phi-winogrande_inverted_option-results" More Information needed
automated-research-group/winogrande
Hugging Face dataset: automated-research-group/winogrande
automated-research-group/winogrande_inverted_option
Dataset Card for "winogrande_inverted_option" More Information needed