Awesome Auto Research Hub Papers · Datasets · Projects
← Back to papers

Can Large Language Models Adequately Perform Symbolic Reasoning Over Time Series?

arXiv 2025 57.3 method

TLDR

Introduces SymbolBench to evaluate LLMs' symbolic reasoning over time series, proposing a framework combining LLMs with genetic programming.

Reasoning

Strengths include a comprehensive benchmark and a novel framework integrating LLMs with genetic programming for symbolic reasoning. Weaknesses are that it focuses narrowly on symbolic reasoning rather than full automated scientific discovery, and the abstract does not detail specific results or limitations.

Read-first score

Read-first score 57.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 34.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,result

Reproducibility 25%
73

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=github

Topical relevance 42%
28.3

Uses existing LLM keyword relevance scores normalized to 0-100. AI scientist,automated scientific discovery,autonomous research agent,automated research,literature review agent,survey generation,automated experimentation,experiment design agent,AI for scientific research,paper writing agent,research automation,scientific discovery agent

Field roles

FrontierMethodology anchorReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 148.

Keyword Scores

AI for scientific research
6
automated scientific discovery
5
research automation
5
scientific discovery agent
5
automated research
4
automated experimentation
3
AI scientist
2
autonomous research agent
2
experiment design agent
2
literature review agent
0
survey generation
0
paper writing agent
0

Deep Analysis

Innovations

  • Introduction of SymbolBench, a comprehensive benchmark for symbolic reasoning over real-world time series, covering multivariate symbolic regression, Boolean network inference, and causal discovery with diverse symbolic forms and complexity.
  • A unified closed-loop framework integrating LLMs with genetic programming, where LLMs serve as both predictors and evaluators.

Methodology

The paper introduces SymbolBench, a benchmark with three symbolic reasoning tasks on real-world time series, and proposes a framework that combines LLMs with genetic programming in a closed-loop system. LLMs are used as predictors and evaluators, and empirical evaluation is conducted on current models.

Key Results

Empirical results reveal key strengths and limitations of current LLMs, emphasizing that domain knowledge, context alignment, and reasoning structure are crucial for improving symbolic reasoning over time series.

Limitations

  • Current LLMs exhibit limitations in symbolic reasoning over time series, necessitating integration with domain knowledge, context alignment, and reasoning structure to achieve adequate performance.

Tags

AI