Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

arXiv 2025 76.9 method

TLDR

Introduces an AI system that autonomously conducts research and writes papers, achieving the first AI-generated workshop paper accepted in peer review.

Reasoning

The paper presents a novel agentic tree-search methodology and demonstrates real-world validation by submitting to a peer-reviewed workshop, with one paper accepted. Strengths include end-to-end automation and empirical evaluation; weaknesses include limited scope (workshop-level) and potential reproducibility concerns.

Read-first score

Read-first score 76.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 92.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Reproducibility 25%
81

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=code,github

Topical relevance 42%
76.7

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Methodology quality 25%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=evaluation,experiment

Field roles

FrontierBridgeMethodology anchorReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 20.

Keyword Scores

AI scientist
10
automated scientific discovery
10
autonomous research agent
9
automated experimentation
9
AI for scientific research
9
paper writing agent
9
scientific discovery agent
9
automated research
8
experiment design agent
8
research automation
8
literature review agent
2
survey generation
1

Deep Analysis

Innovations

  • First entirely AI-generated peer-review-accepted workshop paper
  • Eliminates reliance on human-authored code templates present in v1
  • Generalizes across diverse machine learning domains
  • Progressive agentic tree-search methodology with a dedicated experiment manager agent
  • VLM-based reviewer feedback loop for iterative figure refinement (content and aesthetics)
  • Open-sourced codebase

Methodology

The AI Scientist-v2 is an end-to-end agentic system that iteratively formulates hypotheses, designs and executes experiments, analyzes and visualizes data, and authors manuscripts. It uses a novel progressive agentic tree-search managed by an experiment manager agent, and enhances the AI reviewer with a VLM feedback loop for figure refinement. Evaluation involved submitting three fully autonomous manuscripts to an ICLR workshop.

Key Results

One of the three submitted manuscripts achieved scores exceeding the average human acceptance threshold, marking the first fully AI-generated paper to pass peer review.

Limitations

  • Only one of three submitted manuscripts passed peer review, indicating inconsistent performance
  • Demonstrated only at the workshop level, not full conference publications

Tags

AICLLG