Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer

arXiv 2026 61.2 method, system, application

TLDR

Introduces intervention-centered Auto Research to validate decisions in materials science using held-out transfer, achieving 89.3% preserved orderings.

Reasoning

Strengths include a novel validation method that isolates research decisions, real-world evaluation on ten Matbench endpoints, and clear evidence of a hierarchy. Weaknesses are the domain-specific focus on materials and potential scalability concerns for broader scientific discovery.

Read-first score

Read-first score 61.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 76.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Topical relevance 42%
63.3

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=result,validation

Reproducibility 25%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=artifact,code

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 52.

Keyword Scores

automated research
10
AI scientist
9
AI for scientific research
9
research automation
9
automated scientific discovery
8
autonomous research agent
8
scientific discovery agent
8
automated experimentation
7
experiment design agent
7
literature review agent
1
survey generation
0
paper writing agent
0

Deep Analysis

Innovations

  • Intervention-centered Auto Research that validates research decisions rather than only final pipeline artifacts
  • Independent search over Feature, Model, Representation, and Data axes with inner five-fold feedback and outer holdout matrix
  • Measurement of decision reliability via held-out transfer evidence that the loop never sees
  • Revelation of an information-dependent hierarchy: composition-only tasks support several routes to improvement, structure-informed tasks favor local geometry features and complementary tree ensembles

Methodology

Language-model agents independently search Feature, Model, Representation, and Data axes with inner five-fold cross-validation feedback. Each axis winner is frozen, then an outer holdout matrix compares all alternatives on evidence the loop never sees, across 10 Matbench endpoints with 701 agent-executed attempts.

Key Results

Outer holdout evidence confirmed the selected intervention on 9 of 10 Matbench endpoints, preserved 89.3% of non-tied intervention orderings, and rejected an aggregate Representation gain that inner feedback endorsed. Combining frozen Feature and Model code without further search raised mean outer improvement from 19.0% to 26.3%.

Tags