CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
TLDR
Introduces CausalDS, a benchmark for evaluating causal reasoning in data-science agents using synthetic structural causal models and real-world empirical grounding.
Reasoning
The paper addresses a clear gap between symbolic causal reasoning and data analysis benchmarks, with a novel synthetic generation approach that reduces 'causal parrot' risk. However, the abstract lacks results or comparisons, and the benchmark's effectiveness remains unvalidated.
Read-first score
Read-first score 41.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 26.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 19.