Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

arXiv 2026 43.5 method

TLDR

Introduces DSAgentBench, a benchmark with 275 tasks evaluating agents on end-to-end data-science workflows in real computer environments, showing large capability gaps.

Reasoning

The paper's strength is its realistic benchmark design with deterministic evaluators and extensive model testing. Weaknesses include limited scope to data science rather than broader scientific discovery, and no open-source agent success. The abstract provides clear evidence for real-world evaluation.

Read-first score

Read-first score 43.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 25.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,experiment,result,validation

Reproducibility 25%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,github

Topical relevance 42%
20.8

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 25.

Keyword Scores

AI for scientific research
5
automated experimentation
4
research automation
4
autonomous research agent
3
automated research
3
automated scientific discovery
2
experiment design agent
2
AI scientist
1
scientific discovery agent
1
literature review agent
0
survey generation
0
paper writing agent
0

Tags