Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

MLIPilot: LLM-Driven Auto-Research for Machine-Learned Interatomic Potentials

arXiv 2026 71.9 method

TLDR

MLIPilot uses LLM agents to autonomously optimize machine-learned interatomic potentials via hypothesis proposal, code editing, and HPC job management.

Reasoning

The paper presents a novel framework integrating LLMs with domain-specific constraints for automated MLIP development, with strong empirical evaluation across multiple LLMs and datasets. However, the abstract lacks details on limitations, such as generalizability beyond MLIPs or potential failure modes.

Read-first score

Read-first score 71.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 78.

Methodology quality 25%
100

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=baseline,benchmark,dataset,evaluation,experiment,result,validation

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Topical relevance 42%
65

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,dataset

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 17.

Keyword Scores

automated scientific discovery
9
autonomous research agent
9
automated research
9
scientific discovery agent
9
AI scientist
8
automated experimentation
8
AI for scientific research
8
research automation
8
experiment design agent
7
literature review agent
1
survey generation
1
paper writing agent
1

Deep Analysis

Innovations

  • MLIPilot: an auto-research framework where tool-calling LLM agents propose hypotheses, edit MLIP training code, launch HPC jobs, and accept/revert changes using a fixed, physically constrained scorecard.

Methodology

MLIPilot uses tool-calling LLM agents (GPT-5.5, GPT-4.1, Mistral-24B, Qwen3-32B) to iteratively optimize MACE interatomic potentials by editing code, submitting HPC jobs, and evaluating against a physically constrained scorecard. Evaluation is performed on a QM7-derived molecular dataset with B3LYP/6-31G(d) energies/forces and a Cu EMT periodic dataset with ASE Effective Medium Theory labels.

Key Results

The strongest LLM agents transformed initially constraint-violating baselines into accepted models by autonomously discovering training strategies such as output normalization, loss-function changes, progressive training schedules, and model-capacity adjustments.

Tags

chem-phLG