Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

AI Scientists Are Only as Good as Their Evidence: A Stratified Ablation of Proprietary Data and Reasoning Skills in Drug-Asset Valuation

arXiv 2026 63.3 method

TLDR

Evidence substrate, not reasoning scaffolds, limits AI scientist performance in drug-asset valuation; proprietary data dramatically improves coverage and decision quality.

Reasoning

The paper presents a well-designed ablation study with clear methodology and real-world benchmarks, demonstrating that proprietary evidence is the key bottleneck. However, the findings are domain-specific to drug-asset valuation and may not generalize to other scientific discovery tasks.

Read-first score

Read-first score 63.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 59.

Methodology quality 25%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=ablation,benchmark,dataset,metric,result

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Topical relevance 42%
49.2

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 66.

Keyword Scores

AI scientist
9
autonomous research agent
8
AI for scientific research
8
scientific discovery agent
8
automated scientific discovery
7
automated research
7
research automation
6
literature review agent
2
survey generation
1
automated experimentation
1
experiment design agent
1
paper writing agent
1

Deep Analysis

Innovations

  • Stratified three-arm ablation isolating the effect of proprietary evidence substrate vs. reasoning scaffolds in AI Scientist agents for drug-asset valuation
  • Introduction of completeness-aware decision utility metric: informed decision-quality = decision-quality × gold-coverage
  • Demonstration that proprietary data sets the upper bound of what an AI Scientist can know, with reasoning scaffolds improving calibration and discipline but not removing the factual ceiling

Methodology

Controlled three-arm ablation on a production drug-asset valuation agent across a 13-asset stratified benchmark. Arm A uses a plain web-only LLM analyst; Arm B adds public structured tools, a 14-dimension valuation playbook, verifier, objectivity policy, and red-team; Arm C further adds the proprietary Noah AI corpus of curated pipeline, trial, and deal intelligence. Metrics include tier-in-range accuracy, objectivity scores, gold competitive record recovery, raw decision quality, and the proposed completeness-aware decision utility.

Key Results

Arm B improved tier-in-range accuracy from 0.80 to 0.89 and objectivity from 3.16 to 3.30 but did not lift the factual ceiling; Arm C recovered 0.96 of the gold competitive record versus 0.25/0.30 for A/B, and on completeness-aware decision utility C reached 7.43 versus 1.76/2.57 for A/B, with even a perfect non-proprietary-data report capped at 3.83 by B's coverage.

Limitations

  • Study limited to a single domain (drug-asset valuation) and a specific proprietary dataset (Noah AI corpus)
  • Benchmark size is small (13 assets), potentially limiting generalizability
  • The completeness-aware utility metric depends on gold-coverage estimates that may be dataset-specific
  • Even a perfect non-proprietary-data report is capped by the coverage of public tools, as shown by the 3.83 ceiling

Tags

AI