Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

Grounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physics

arXiv 2026 70.5 method

TLDR

An LLM pipeline autonomously conducts literature review, reproduces experiments, performs novel computations, and writes a manuscript in condensed-matter physics with fault tolerance.

Reasoning

Strengths: novel fault-tolerant pipeline with grounding in literature, real-world application in physics. Weaknesses: limited to one domain, requires human intervention at reproduction failures, scalability unclear.

Read-first score

Read-first score 70.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 94.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=ablation,baseline,result

Topical relevance 42%
78.3

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=checkpoint

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 44.

Keyword Scores

automated scientific discovery
9
autonomous research agent
9
research automation
9
scientific discovery agent
9
AI scientist
8
automated research
8
literature review agent
8
AI for scientific research
8
paper writing agent
8
automated experimentation
7
experiment design agent
6
survey generation
5

Deep Analysis

Innovations

  • End-to-end autonomous research pipeline from a literature corpus to a publication-grade manuscript in computational physics, with literature grounding throughout.
  • Fault tolerance via fresh-context isolation, distributed grounding, and adversarial review across 47 sessions sharing only on-disk state.
  • Calibration by reproducing published references to ground methodology, preventing hallucination.
  • Structurally enforced numerical confrontation at calibration checkpoints as the operative grounding mechanism, isolated via paired failure-mode ablations.
  • Bounded human intervention limited to operational knowledge curation at reproduction failures, not scientific direction.

Methodology

The pipeline ingests a corpus of 11,083 condensed-matter physics arXiv papers, autonomously maps the corpus to conceive a research direction, calibrates by reproducing published references, conducts novel first-principles computations, and writes a manuscript. It operates across six phases in 47 fresh-context sessions that share only on-disk state, with 2,162 literature-consultation events. Fault tolerance is achieved through redundancy: fresh-context isolation, distributed grounding, and adversarial review. Pre- and post-pilot stages are fully autonomous; the pilot stage requires human intervention only when reproduction fails. Two ablations (pre-architecture baseline and no-pilot) isolate the calibration-checkpoint grounding mechanism.

Key Results

The pipeline produced a publication-grade manuscript with three substantive physics findings on altermagnetic piezomagnetism, grounded in literature throughout, and the fault-tolerant design with calibration checkpoints prevented hallucination while quantifying the intervention pattern.

Limitations

  • The pilot stage requires bounded human intervention at reproduction failures, so the pipeline is not fully autonomous end-to-end.
  • Demonstrated only in a single computational physics subdomain (altermagnetic piezomagnetism) using a specific corpus of recent arXiv papers.
  • Relies on the availability of a large, recent literature corpus for calibration; generalization to domains without such anchors is not shown.

Tags