Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

arXiv 2026 50.7 method

TLDR

Proposes using adversarial fast-moving real-world domains (F1, MTG) as benchmarks for AI scientist capabilities, finding models produce plausible ideas but few match expert solutions.

Reasoning

The paper introduces a practical benchmarking framework using real-world expert outputs, which addresses limitations of synthetic tasks and retrospective targets. Its strengths include novel domains and concrete evaluation metrics, but weaknesses include limited domain scope and potential subjectivity in defining ground truth for F1 innovations.

Read-first score

Read-first score 50.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 50.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,result

Topical relevance 42%
41.7

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 40.

Keyword Scores

AI scientist
10
scientific discovery agent
9
automated scientific discovery
8
AI for scientific research
8
automated research
5
autonomous research agent
4
research automation
4
automated experimentation
1
experiment design agent
1
literature review agent
0
survey generation
0
paper writing agent
0

Deep Analysis

Innovations

  • Proposing adversarial, fast-moving real-world domains as test beds for benchmarking AI scientist capabilities
  • Instantiation in Formula 1 car design ideation for the 2026 season with real pre-season innovations as ground truth
  • Instantiation in Magic: The Gathering deck building with Pro Tour decklists as evaluation targets

Methodology

The framework evaluates AI scientist capabilities by having models generate novel ideas in two complex, adversarial, fast-moving domains: Formula 1 (car design concepts for 2026) and Magic: The Gathering (decks from a recently updated card pool). Outputs are compared against independently produced expert ground truth (real F1 innovations and 19 Pro Tour decklists) to measure alignment.

Key Results

In F1, GPT-5.2 matched 10 of 40 real innovations across 166 ideas; in MTG, the best deck recovered 5 of 7 new-set cards from a top Pro Tour deck, and model card selections correlated with Pro Tour adoption (Spearman ρ=0.74, p=0.0003).

Tags