DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
TLDR
Introduces DISCOVERYWORLD, a virtual environment with 120 tasks for benchmarking automated scientific discovery agents.
Reasoning
Strengths include a novel, comprehensive virtual environment with diverse tasks and automatic metrics; weaknesses are that it is simulated rather than real-world, and baseline agents struggle, indicating difficulty but also potential limitations in generalizability.
Read-first score
Read-first score 77.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 75.
Field roles
Rank sensitivity
Stability: volatile; rank range: 28.
Keyword Scores
Deep Analysis
Innovations
- First virtual environment for developing and benchmarking agents on complete cycles of novel scientific discovery.
- 120 diverse challenge tasks across 8 topics with 3 difficulty levels and parametric variations, covering radioisotope dating, rocket science, proteomics, etc.
- Three automatic evaluation metrics: task completion, task-relevant actions, and discovered explanatory knowledge.
- Demonstration that strong baseline agents struggle, highlighting the environment's ability to capture novel discovery challenges.
Methodology
DISCOVERYWORLD is a simulated text-based environment (with optional 2D visual overlay) containing 120 tasks across 8 scientific topics. Each task requires an agent to form hypotheses, design and run experiments, analyze results, and act on conclusions. Performance is evaluated using three automatic metrics: task completion, task-relevant actions taken, and discovered explanatory knowledge.
Key Results
Strong baseline agents that perform well in prior environments struggle on most DISCOVERYWORLD tasks, indicating the environment captures novel challenges of scientific discovery.