Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities
TLDR
Proposes using adversarial fast-moving real-world domains (F1, MTG) as benchmarks for AI scientist capabilities, finding models produce plausible ideas but few match expert solutions.
Reasoning
The paper introduces a practical benchmarking framework using real-world expert outputs, which addresses limitations of synthetic tasks and retrospective targets. Its strengths include novel domains and concrete evaluation metrics, but weaknesses include limited domain scope and potential subjectivity in defining ground truth for F1 innovations.
Read-first score
Read-first score 50.7, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 50.
Field roles
Rank sensitivity
Stability: volatile; rank range: 40.
Keyword Scores
Deep Analysis
Innovations
- Proposing adversarial, fast-moving real-world domains as test beds for benchmarking AI scientist capabilities
- Instantiation in Formula 1 car design ideation for the 2026 season with real pre-season innovations as ground truth
- Instantiation in Magic: The Gathering deck building with Pro Tour decklists as evaluation targets
Methodology
The framework evaluates AI scientist capabilities by having models generate novel ideas in two complex, adversarial, fast-moving domains: Formula 1 (car design concepts for 2026) and Magic: The Gathering (decks from a recently updated card pool). Outputs are compared against independently produced expert ground truth (real F1 innovations and 19 Pro Tour decklists) to measure alignment.
Key Results
In F1, GPT-5.2 matched 10 of 40 real innovations across 166 ideas; in MTG, the best deck recovered 5 of 7 new-set cards from a top Pro Tour deck, and model card selections correlated with Pro Tour adoption (Spearman ρ=0.74, p=0.0003).