Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer
TLDR
Introduces intervention-centered Auto Research to validate decisions in materials science using held-out transfer, achieving 89.3% preserved orderings.
Reasoning
Strengths include a novel validation method that isolates research decisions, real-world evaluation on ten Matbench endpoints, and clear evidence of a hierarchy. Weaknesses are the domain-specific focus on materials and potential scalability concerns for broader scientific discovery.
Read-first score
Read-first score 61.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 76.
Field roles
Rank sensitivity
Stability: volatile; rank range: 52.
Keyword Scores
Deep Analysis
Innovations
- Intervention-centered Auto Research that validates research decisions rather than only final pipeline artifacts
- Independent search over Feature, Model, Representation, and Data axes with inner five-fold feedback and outer holdout matrix
- Measurement of decision reliability via held-out transfer evidence that the loop never sees
- Revelation of an information-dependent hierarchy: composition-only tasks support several routes to improvement, structure-informed tasks favor local geometry features and complementary tree ensembles
Methodology
Language-model agents independently search Feature, Model, Representation, and Data axes with inner five-fold cross-validation feedback. Each axis winner is frozen, then an outer holdout matrix compares all alternatives on evidence the loop never sees, across 10 Matbench endpoints with 701 agent-executed attempts.
Key Results
Outer holdout evidence confirmed the selected intervention on 9 of 10 Matbench endpoints, preserved 89.3% of non-tied intervention orderings, and rejected an aggregate Representation gain that inner feedback endorsed. Combining frozen Feature and Model code without further search raised mean outer improvement from 19.0% to 26.3%.