Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR

arXiv 2026 54.7 method, benchmark

TLDR

Trains search agents to reason over scientific papers using RLVR, releasing a biomedical corpus and QA dataset.

Reasoning

Strengths include novel RLVR training for scientific search, large corpus, and dataset. Weaknesses are limited to biomedical factoid QA and not full scientific discovery.

Read-first score

Read-first score 54.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 43.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,baseline,benchmark,dataset

Reproducibility 25%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,dataset

Topical relevance 42%
35.8

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

FrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 70.

Keyword Scores

literature review agent
8
AI for scientific research
7
automated research
6
research automation
6
autonomous research agent
5
scientific discovery agent
4
AI scientist
3
automated scientific discovery
2
survey generation
2
automated experimentation
0
experiment design agent
0
paper writing agent
0

Deep Analysis

Innovations

  • Training search agents to search and reason over scientific papers using RLVR, targeting technical QA in science/engineering/medicine.
  • Release of a 16M biomedical paper abstract search corpus and a 60k-sample factoid QA dataset (PaperSearchQA) with benchmarks.
  • Demonstration that RLVR-trained agents outperform non-RL retrieval baselines.
  • Observation of agent behaviors like planning, reasoning, and self-verification.
  • Scalable data creation methods extendable to other scientific domains.

Methodology

The authors construct a search corpus of 16 million biomedical paper abstracts and a factoid QA dataset (PaperSearchQA) with 60k samples answerable from the corpus. They train search agents using reinforcement learning with verifiable rewards (RLVR) in this environment, comparing against non-RL retrieval baselines. They also perform quantitative analysis of agent behaviors.

Key Results

RLVR-trained search agents outperform non-RL retrieval baselines, and agents exhibit planning, reasoning, and self-verification behaviors.

Tags

LGAICLIR