Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Speech World Model: Causal State-Action Planning with Explicit Reasoning for Speech

arXiv 25.12 2025 33.5 method

TLDR

Proposes a modular speech world model with causal graph for explicit reasoning and interpretability under partial supervision.

Reasoning

Strengths include a novel modular causal graph approach inspired by cognitive science, enabling counterfactual interventions and interpretability. Weaknesses are the lack of empirical results or real-world evaluation in the abstract, and reliance on future open-sourcing without evidence of performance.

Read-first score

Read-first score 33.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 19.

Recency 6%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 18%
70

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,experiment,result

Reproducibility 18%
46

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code,open source

Topical relevance 29%
27.1

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Citation impact 18%
0

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. cited_by_count=0

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 184.

Keyword Scores

world model
10
world dynamics prediction
8
generative world model
1
world simulator
0
interactive world model
0
video world model
0
model-based reinforcement learning world model
0

Deep Analysis

Innovations

  • Explicit reasoning over speech states and actions with modular and transparent decisions
  • Factorization of speech understanding into four modules communicating through a causal graph
  • Cognitive state search space guided by posterior traces
  • Instruction-tuned language model for causal analysis and user-facing response
  • Counterfactual interventions and interpretability under partial supervision
  • First graph-based modular speech model for explicit reasoning

Methodology

The system adopts a world model perspective, learning forward dynamics over latent states. It factorizes speech understanding into four modules that interact via a causal graph, establishing a cognitive state search space. An instruction-tuned language model uses posterior traces from this space to generate causal analysis and responses.

Key Results

The abstract does not report quantitative experimental results. It claims the model enables counterfactual interventions and interpretability under partial supervision, and that it is the first graph-based modular speech model for explicit reasoning.

Tags