Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence

arXiv 2026 62.9 method

TLDR

Apodex Discovery introduces benchmarks and environments for discoverative AI, with a heavy-duty solver and HDS6 evaluation, surpassing baselines in AAV capsid design and drug repurposing.

Reasoning

Strengths include a concrete framework, real-world problem selection, independent evaluation metrics, and empirical results in two domains. Weaknesses are that the abstract is truncated, with limited details on baselines and reproducibility, and the new terminology may obscure comparisons to existing agent frameworks.

Read-first score

Read-first score 62.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 65.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=ablation,baseline,benchmark,evaluation,metric

Topical relevance 42%
54.2

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=artifact

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 39.

Keyword Scores

automated scientific discovery
9
AI for scientific research
9
scientific discovery agent
9
autonomous research agent
7
automated experimentation
7
AI scientist
6
automated research
6
research automation
6
experiment design agent
5
survey generation
1
literature review agent
0
paper writing agent
0

Deep Analysis

Innovations

  • Heavy-duty solver architecture: a foundation model with harness, tools, and control policies for extended, stateful, verifiable investigations
  • Problem-scouting process: systematically surveyed 561 industries across 16 sectors to assemble 423 high-value real-world problems, with 20 selected for initial release
  • Common environment-task-episode abstraction providing data, tools, constraints, feedback, trajectory recording, and verification of intermediate and final artifacts
  • HDS6 evaluation metric: independently assesses Tools, Repair, Alternatives, Coherence, Evidence, and Scope beyond final-task success
  • TRACES episode interface enabling attribution of performance differences to specific solver components via controlled ablations

Methodology

Apodex Discovery introduces a framework for discoverative AI built around a heavy-duty solver that combines a foundation model, harness, tools, and control policies. It defines a standardized environment-task-episode abstraction that supplies data, tools, constraints, feedback, and verification of intermediate artifacts and final submissions. Evaluation is performed using the HDS6 rubric, which scores Tools, Repair, Alternatives, Coherence, Evidence, and Scope independently of final-task success.

Key Results

In AAV capsid design, Apodex surpassed the published state of the art by 7% across viability, tropism, structure prediction, and generative design. In drug repurposing and reformulation, a biomedical environment improved GPT-5.5 and GPT-5.6-sol mean normalized prediction scores by 2.5 and 7.6 points over closed-book baselines, and controlled ablations using the TRACES interface isolated performance contributions to specific solver components.

Tags