AI Scientists Are Only as Good as Their Evidence: A Stratified Ablation of Proprietary Data and Reasoning Skills in Drug-Asset Valuation
TLDR
Evidence substrate, not reasoning scaffolds, limits AI scientist performance in drug-asset valuation; proprietary data dramatically improves coverage and decision quality.
Reasoning
The paper presents a well-designed ablation study with clear methodology and real-world benchmarks, demonstrating that proprietary evidence is the key bottleneck. However, the findings are domain-specific to drug-asset valuation and may not generalize to other scientific discovery tasks.
Read-first score
Read-first score 63.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 59.
Field roles
Rank sensitivity
Stability: volatile; rank range: 66.
Keyword Scores
Deep Analysis
Innovations
- Stratified three-arm ablation isolating the effect of proprietary evidence substrate vs. reasoning scaffolds in AI Scientist agents for drug-asset valuation
- Introduction of completeness-aware decision utility metric: informed decision-quality = decision-quality × gold-coverage
- Demonstration that proprietary data sets the upper bound of what an AI Scientist can know, with reasoning scaffolds improving calibration and discipline but not removing the factual ceiling
Methodology
Controlled three-arm ablation on a production drug-asset valuation agent across a 13-asset stratified benchmark. Arm A uses a plain web-only LLM analyst; Arm B adds public structured tools, a 14-dimension valuation playbook, verifier, objectivity policy, and red-team; Arm C further adds the proprietary Noah AI corpus of curated pipeline, trial, and deal intelligence. Metrics include tier-in-range accuracy, objectivity scores, gold competitive record recovery, raw decision quality, and the proposed completeness-aware decision utility.
Key Results
Arm B improved tier-in-range accuracy from 0.80 to 0.89 and objectivity from 3.16 to 3.30 but did not lift the factual ceiling; Arm C recovered 0.96 of the gold competitive record versus 0.25/0.30 for A/B, and on completeness-aware decision utility C reached 7.43 versus 1.76/2.57 for A/B, with even a perfect non-proprietary-data report capped at 3.83 by B's coverage.
Limitations
- Study limited to a single domain (drug-asset valuation) and a specific proprietary dataset (Noah AI corpus)
- Benchmark size is small (13 assets), potentially limiting generalizability
- The completeness-aware utility metric depends on gold-coverage estimates that may be dataset-specific
- Even a perfect non-proprietary-data report is capped by the coverage of public tools, as shown by the 3.83 ceiling