The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
TLDR
Introduces an AI system that autonomously conducts research and writes papers, achieving the first AI-generated workshop paper accepted in peer review.
Reasoning
The paper presents a novel agentic tree-search methodology and demonstrates real-world validation by submitting to a peer-reviewed workshop, with one paper accepted. Strengths include end-to-end automation and empirical evaluation; weaknesses include limited scope (workshop-level) and potential reproducibility concerns.
Read-first score
Read-first score 76.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 92.
Field roles
Rank sensitivity
Stability: volatile; rank range: 20.
Keyword Scores
Deep Analysis
Innovations
- First entirely AI-generated peer-review-accepted workshop paper
- Eliminates reliance on human-authored code templates present in v1
- Generalizes across diverse machine learning domains
- Progressive agentic tree-search methodology with a dedicated experiment manager agent
- VLM-based reviewer feedback loop for iterative figure refinement (content and aesthetics)
- Open-sourced codebase
Methodology
The AI Scientist-v2 is an end-to-end agentic system that iteratively formulates hypotheses, designs and executes experiments, analyzes and visualizes data, and authors manuscripts. It uses a novel progressive agentic tree-search managed by an experiment manager agent, and enhances the AI reviewer with a VLM feedback loop for figure refinement. Evaluation involved submitting three fully autonomous manuscripts to an ICLR workshop.
Key Results
One of the three submitted manuscripts achieved scores exceeding the average human acceptance threshold, marking the first fully AI-generated paper to pass peer review.
Limitations
- Only one of three submitted manuscripts passed peer review, indicating inconsistent performance
- Demonstrated only at the workshop level, not full conference publications