Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

MegaScience: Pushing the Frontiers of Post-Training Datasets for Science Reasoning

arXiv 2025 56 benchmark

TLDR

Introduces open-source scientific reasoning datasets (TextbookReasoning and MegaScience) to improve AI scientist training, outperforming existing datasets.

Reasoning

Strengths include large-scale, high-quality, verifiable datasets from textbooks and systematic ablation studies; weaknesses are limited to post-training datasets without addressing full automation or real-world experimental validation.

Read-first score

Read-first score 56, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 41.

Methodology quality 25%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=ablation,benchmark,dataset,evaluation,experiment,metric

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Topical relevance 42%
34.2

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

FrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 87.

Keyword Scores

AI for scientific research
8
AI scientist
7
scientific discovery agent
5
automated scientific discovery
4
research automation
4
autonomous research agent
3
automated research
3
automated experimentation
2
experiment design agent
2
literature review agent
1
survey generation
1
paper writing agent
1

Deep Analysis

Innovations

  • TextbookReasoning dataset: 650k reasoning questions with truthful reference answers extracted from 12k university-level scientific textbooks across 7 disciplines
  • MegaScience dataset: a 1.25M-instance mixture of high-quality open-source scientific datasets, curated via systematic ablation studies to identify optimal subsets
  • Comprehensive evaluation system with 15 benchmarks and robust answer extraction strategies for accurate metrics

Methodology

They extracted TextbookReasoning from 12k textbooks, then built MegaScience by mixing open-source scientific datasets, using ablation studies to select optimal subsets. They trained Llama3.1, Qwen2.5, Qwen3 base models on MegaScience and evaluated on 15 benchmarks with careful answer extraction.

Key Results

MegaScience-trained models achieve superior performance and training efficiency with shorter responses, significantly outperforming official instruct models on average, with larger models benefiting more, indicating scaling benefits for scientific tuning.

Tags

CLAILG