Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents

arXiv 2024 77.5 method

TLDR

Introduces DISCOVERYWORLD, a virtual environment with 120 tasks for benchmarking automated scientific discovery agents.

Reasoning

Strengths include a novel, comprehensive virtual environment with diverse tasks and automatic metrics; weaknesses are that it is simulated rather than real-world, and baseline agents struggle, indicating difficulty but also potential limitations in generalizability.

Read-first score

Read-first score 77.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 75.

Methodology quality 25%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,benchmark,evaluation,experiment,metric,result

Reproducibility 25%
81

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=code,github

Recency 8%
75.1

Uses a gentle age decay so recent papers surface without erasing older foundations. 2024

Topical relevance 42%
62.5

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

Methodology anchorReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 28.

Keyword Scores

automated scientific discovery
10
scientific discovery agent
10
automated experimentation
9
autonomous research agent
8
experiment design agent
8
automated research
7
AI for scientific research
7
research automation
7
AI scientist
6
literature review agent
1
survey generation
1
paper writing agent
1

Deep Analysis

Innovations

  • First virtual environment for developing and benchmarking agents on complete cycles of novel scientific discovery.
  • 120 diverse challenge tasks across 8 topics with 3 difficulty levels and parametric variations, covering radioisotope dating, rocket science, proteomics, etc.
  • Three automatic evaluation metrics: task completion, task-relevant actions, and discovered explanatory knowledge.
  • Demonstration that strong baseline agents struggle, highlighting the environment's ability to capture novel discovery challenges.

Methodology

DISCOVERYWORLD is a simulated text-based environment (with optional 2D visual overlay) containing 120 tasks across 8 scientific topics. Each task requires an agent to form hypotheses, design and run experiments, analyze results, and act on conclusions. Performance is evaluated using three automatic metrics: task completion, task-relevant actions taken, and discovered explanatory knowledge.

Key Results

Strong baseline agents that perform well in prior environments struggle on most DISCOVERYWORLD tasks, indicating the environment captures novel challenges of scientific discovery.

Tags

AICL