Can Large Language Models Adequately Perform Symbolic Reasoning Over Time Series?
TLDR
Introduces SymbolBench to evaluate LLMs' symbolic reasoning over time series, proposing a framework combining LLMs with genetic programming.
Reasoning
Strengths include a comprehensive benchmark and a novel framework integrating LLMs with genetic programming for symbolic reasoning. Weaknesses are that it focuses narrowly on symbolic reasoning rather than full automated scientific discovery, and the abstract does not detail specific results or limitations.
Read-first score
Read-first score 57.3, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 34.
Field roles
Rank sensitivity
Stability: volatile; rank range: 148.
Keyword Scores
Deep Analysis
Innovations
- Introduction of SymbolBench, a comprehensive benchmark for symbolic reasoning over real-world time series, covering multivariate symbolic regression, Boolean network inference, and causal discovery with diverse symbolic forms and complexity.
- A unified closed-loop framework integrating LLMs with genetic programming, where LLMs serve as both predictors and evaluators.
Methodology
The paper introduces SymbolBench, a benchmark with three symbolic reasoning tasks on real-world time series, and proposes a framework that combines LLMs with genetic programming in a closed-loop system. LLMs are used as predictors and evaluators, and empirical evaluation is conducted on current models.
Key Results
Empirical results reveal key strengths and limitations of current LLMs, emphasizing that domain knowledge, context alignment, and reasoning structure are crucial for improving symbolic reasoning over time series.
Limitations
- Current LLMs exhibit limitations in symbolic reasoning over time series, necessitating integration with domain knowledge, context alignment, and reasoning structure to achieve adequate performance.