MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning
TLDR
Introduces open-source scientific reasoning datasets (TextbookReasoning and MegaScience) to improve AI scientist training, outperforming existing datasets.
Reasoning
Strengths include large-scale, high-quality, verifiable datasets from textbooks and systematic ablation studies; weaknesses are limited to post-training datasets without addressing full automation or real-world experimental validation.
Read-first score
Read-first score 56, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 41.
Field roles
Rank sensitivity
Stability: volatile; rank range: 87.
Keyword Scores
Deep Analysis
Innovations
- TextbookReasoning dataset: 650k reasoning questions with truthful reference answers extracted from 12k university-level scientific textbooks across 7 disciplines
- MegaScience dataset: a 1.25M-instance mixture of high-quality open-source scientific datasets, curated via systematic ablation studies to identify optimal subsets
- Comprehensive evaluation system with 15 benchmarks and robust answer extraction strategies for accurate metrics
Methodology
They extracted TextbookReasoning from 12k textbooks, then built MegaScience by mixing open-source scientific datasets, using ablation studies to select optimal subsets. They trained Llama3.1, Qwen2.5, Qwen3 base models on MegaScience and evaluated on 15 benchmarks with careful answer extraction.
Key Results
MegaScience-trained models achieve superior performance and training efficiency with shorter responses, significantly outperforming official instruct models on average, with larger models benefiting more, indicating scaling benefits for scientific tuning.