DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
TLDR
Introduces DSAgentBench, a benchmark with 275 tasks evaluating agents on end-to-end data-science workflows in real computer environments, showing large capability gaps.
Reasoning
The paper's strength is its realistic benchmark design with deterministic evaluators and extensive model testing. Weaknesses include limited scope to data science rather than broader scientific discovery, and no open-source agent success. The abstract provides clear evidence for real-world evaluation.
Read-first score
Read-first score 43.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 25.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 25.