Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

HARPA: A Testability-Driven, Literature-Grounded Framework for Research Ideation

arXiv 2025 54 method

TLDR

HARPA is a framework for generating testable, literature-grounded research hypotheses, showing gains in feasibility and groundedness over baselines.

Reasoning

The paper presents a novel ideation framework that integrates literature mining and hypothesis design, with strong empirical validation including comparisons to a baseline and an ASD agent. Strengths include clear methodology and significant improvements in feasibility and groundedness; weaknesses may include limited scope of evaluation (only one ASD agent) and potential over-reliance on LLMs.

Read-first score

Read-first score 54, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 57.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,evaluation,experiment

Topical relevance 42%
47.5

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 35.

Keyword Scores

automated scientific discovery
8
AI for scientific research
7
research automation
7
AI scientist
6
automated research
6
scientific discovery agent
6
autonomous research agent
5
literature review agent
4
experiment design agent
3
paper writing agent
3
automated experimentation
2
survey generation
0

Deep Analysis

Innovations

  • Testability-driven, literature-grounded framework for hypothesis generation that ensures hypotheses are both testable and grounded in scientific literature.
  • Human-inspired ideation workflow: literature mining for emerging trends, exploration of hypothesis design spaces, and convergence on precise testable hypotheses by pinpointing research gaps and justifying design choices.
  • Adaptive reward model that learns from prior experimental outcomes to score new hypotheses, enabling continuous refinement of hypothesis quality.
  • Significant improvements in feasibility and groundedness over a strong baseline AI-researcher, with corresponding gains in execution success when used with an ASD agent.

Methodology

HARPA mines scientific literature to identify emerging trends, explores hypothesis design spaces, and converges on testable hypotheses by pinpointing research gaps and justifying design choices. It learns a reward model from prior experimental outcomes to score hypotheses. Evaluations compare HARPA-generated proposals to a baseline AI-researcher on qualitative dimensions (specificity, novelty, overall quality, feasibility, groundedness) using a 10-point Likert scale, and test execution success with the CodeScientist ASD agent.

Key Results

HARPA proposals matched baseline on specificity, novelty, and overall quality but significantly improved feasibility (+0.78, p<0.05) and groundedness (+0.85, p<0.01). When used with CodeScientist, HARPA achieved 20 successful vs 11 baseline executions out of 40, and fewer failures (16 vs 21). The learned reward model yielded ~28% absolute gain over untrained baseline scorer.

Tags

AICL