HARPA: A Testability-Driven, Literature-Grounded Framework for Research Ideation
TLDR
HARPA is a framework for generating testable, literature-grounded research hypotheses, showing gains in feasibility and groundedness over baselines.
Reasoning
The paper presents a novel ideation framework that integrates literature mining and hypothesis design, with strong empirical validation including comparisons to a baseline and an ASD agent. Strengths include clear methodology and significant improvements in feasibility and groundedness; weaknesses may include limited scope of evaluation (only one ASD agent) and potential over-reliance on LLMs.
Read-first score
Read-first score 54, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 57.
Field roles
Rank sensitivity
Stability: volatile; rank range: 35.
Keyword Scores
Deep Analysis
Innovations
- Testability-driven, literature-grounded framework for hypothesis generation that ensures hypotheses are both testable and grounded in scientific literature.
- Human-inspired ideation workflow: literature mining for emerging trends, exploration of hypothesis design spaces, and convergence on precise testable hypotheses by pinpointing research gaps and justifying design choices.
- Adaptive reward model that learns from prior experimental outcomes to score new hypotheses, enabling continuous refinement of hypothesis quality.
- Significant improvements in feasibility and groundedness over a strong baseline AI-researcher, with corresponding gains in execution success when used with an ASD agent.
Methodology
HARPA mines scientific literature to identify emerging trends, explores hypothesis design spaces, and converges on testable hypotheses by pinpointing research gaps and justifying design choices. It learns a reward model from prior experimental outcomes to score hypotheses. Evaluations compare HARPA-generated proposals to a baseline AI-researcher on qualitative dimensions (specificity, novelty, overall quality, feasibility, groundedness) using a 10-point Likert scale, and test execution success with the CodeScientist ASD agent.
Key Results
HARPA proposals matched baseline on specificity, novelty, and overall quality but significantly improved feasibility (+0.78, p<0.05) and groundedness (+0.85, p<0.01). When used with CodeScientist, HARPA achieved 20 successful vs 11 baseline executions out of 40, and fewer failures (16 vs 21). The learned reward model yielded ~28% absolute gain over untrained baseline scorer.