When Should an AI Scientist Stop? Verifiable Experiment Steering and Refusal for Autonomous Discovery
TLDR
CARTOGRAPH is a verification layer for AI scientists that uses experiment steering, ambiguity closure, and refusal to decide when to stop, validated across multiple testbeds and real-world audits.
Reasoning
The paper introduces a theoretically grounded verification layer with strong empirical results across diverse testbeds, including a retrospective audit of real A-Lab claims. However, the local linear-Gaussian assumption may limit generalization to non-linear settings, and the abstract does not detail comparisons to other stopping criteria.
Read-first score
Read-first score 72.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 70.
Field roles
Rank sensitivity
Stability: volatile; rank range: 37.
Keyword Scores
Deep Analysis
Innovations
- CARTOGRAPH verification layer coupling unresolved-subspace experiment steering (select), explicit ambiguity closure (resolve), and residual-based library inadequacy detection (refuse)
- Exact unresolved A-optimal rule (CARTOGRAPH-A) derived under a local linear-Gaussian bridge, with raw unresolved projection shown as isotropic Fisher-information trace
- Closed-form expected information gain and Box-Hill reinterpreted as local comparators rather than global equivalents
- Refusal mechanism that tentatively identifies out-of-library mechanisms and then revokes them when residuals expose structural misfit
- Retrospective audit of A-Lab autonomous materials claims, flagging all inconclusive claims while passing confirmed ones
Methodology
The paper introduces CARTOGRAPH, a verification layer that selects experiments via unresolved-subspace steering, resolves ambiguity explicitly, and refuses when residual-based library inadequacy is detected. Under a local linear-Gaussian bridge, it derives CARTOGRAPH-A as the exact A-optimal rule and compares it to raw projection and closed-form EIG/Box-Hill. Evaluation spans five testbeds including a structured cascade, pharmacokinetic mechanism identification, filtered EPA settings, and a retrospective audit of 40 positive claims from the A-Lab autonomous materials system.
Key Results
CARTOGRAPH-A beats raw projection 129W/0T/15L at d=8 (p ≈ 10⁻²¹) in a replicated structured cascade; it identifies then revokes three out-of-library pharmacokinetic mechanisms while keeping an in-library control; in the A-Lab audit, the refuse guard flags all 4 claims later marked inconclusive and passes 32/36 confirmed claims.
Limitations
- Derivations rely on a local linear-Gaussian bridge assumption, which may not hold globally
- In low-dimensional pharmacokinetic and filtered EPA settings, near-ties against disagreement are observed, indicating limited advantage
- The refusal mechanism initially misidentifies out-of-library mechanisms before revoking them, so real-time reliability depends on residual detection
- Retrospective audit is limited to a single published system (A-Lab), and generalizability to other autonomous discovery systems is not demonstrated
- Residual-based library inadequacy detection may fail to flag structural misfit if residuals do not clearly expose it