Scaling Scientific Discovery Environments for Turn-Level Agentic RL
TLDR
SciDisco: a scalable framework for training scientific discovery agents using process-verifiable environments and turn-level reinforcement learning.
Reasoning
The paper introduces a well-structured framework (SciDisco) with clear components (SciThèque, DAG trajectory synthesis, DiscoPO) and achieves SOTA on benchmarks, demonstrating strong methodology. However, the abstract lacks details on the specific datasets and real-world applicability, and the scope is limited to hypothesis-driven data analysis rather than full scientific discovery.
Read-first score
Read-first score 58.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 60.
Field roles
Rank sensitivity
Stability: volatile; rank range: 21.
Keyword Scores
Deep Analysis
Innovations
- SciDisco framework for training scientific discovery agents in process-verifiable environments
- SciThèque: compilation of hypotheses, datasets, hidden evidence graphs, and verifiers into task environments with progress checks
- DAG-grounded trajectory synthesis to construct verifier-filtered multi-turn demonstrations
- DiscoPO: turn-level credit assignment using environment as training signal, rewarding actions that produce verifiable analytical evidence
Methodology
The paper proposes SciDisco, a framework that uses SciThèque to create process-verifiable environments with hidden evidence graphs and verifiers. DAG-grounded trajectory synthesis generates multi-turn demonstrations filtered by verifiers, and DiscoPO assigns turn-level credit for actions that yield verifiable evidence. A 14B model is trained and evaluated on hypothesis-driven scientific data analysis benchmarks.
Key Results
SciDisco-14B achieves state-of-the-art performance on hypothesis-driven scientific data analysis benchmarks.