Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

CausalDS: Benchmarking Causal Reasoning in Data-Science Agents

arXiv 2026 41.9 method

TLDR

Introduces CausalDS, a benchmark for evaluating causal reasoning in data-science agents using synthetic structural causal models and real-world empirical grounding.

Reasoning

The paper addresses a clear gap between symbolic causal reasoning and data analysis benchmarks, with a novel synthetic generation approach that reduces 'causal parrot' risk. However, the abstract lacks results or comparisons, and the benchmark's effectiveness remains unvalidated.

Read-first score

Read-first score 41.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 26.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,benchmark,dataset,evaluation

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Topical relevance 42%
21.7

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 19.

Keyword Scores

autonomous research agent
5
AI for scientific research
4
automated scientific discovery
3
automated experimentation
3
scientific discovery agent
3
AI scientist
2
automated research
2
experiment design agent
2
research automation
2
literature review agent
0
survey generation
0
paper writing agent
0

Tags