Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama

arXiv 2026 57.9 benchmark, method, application

TLDR

A benchmark and latent world model for long-horizon audio drama, showing frontier LLMs saturate while N-VSSM maintains high consistency across 200 episodes.

Reasoning

Strengths include a novel benchmark with cross-lingual evaluation and a latent world model that outperforms frontier LLMs on long-arc consistency with lower compute, validated by a human study. Weaknesses are domain specificity to audio drama and lack of generalization evidence to other modalities.

Read-first score

Read-first score 57.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 34.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
95.7

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.95681112

Methodology quality 18%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,metric

Topical relevance 29%
48.6

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 404.

Keyword Scores

world model
9
generative world model
8
world dynamics prediction
7
interactive world model
6
world simulator
3
model-based reinforcement learning world model
1
video world model
0

Deep Analysis

Innovations

  • NarrativeWorldBench: an open benchmark of nine narrative-structure metrics evaluated across horizons h in {10, 20, 50, 100, 200} with cross-lingual evaluation across four Indic languages (Hindi, Tamil, Telugu, Marathi).
  • N-VSSM: a Narrative Variational State-Space Model that maintains a structured 256-dimensional latent world state over more than 200 episodes via a Mamba-2 backbone with an event-conditioned posterior and an 8B decoder.
  • A learned Cultural Transfer Function that lifts cross-language fidelity by +0.20 to +0.23 Likert points.

Methodology

The paper benchmarks 21 models across classical, fine-tuned, open-frontier, closed-frontier, and reasoning tiers on a uniform set of structural narrative metrics. It introduces NarrativeWorldBench with nine metrics and cross-lingual evaluation, and proposes N-VSSM, a latent world model using a Mamba-2 backbone and event-conditioned posterior. A within-subjects writer study with 12 professional authors and 240 trials compares N-VSSM against Claude Opus 4.5.

Key Results

N-VSSM achieves plot-beat F1 = 0.84 across all horizons at 4x lower compute than the closed-frontier band. The Cultural Transfer Function improves cross-language fidelity by +0.20 to +0.23 Likert points. In the writer study, N-VSSM is preferred over Claude Opus 4.5 on long-arc consistency 71% of the time and rated +1.3 Likert points higher on controllability.

Limitations

  • Benchmark only covers four Indic languages, limiting cross-lingual generality.
  • Writer study uses only 12 professional authors, which may not be representative.
  • Model's performance is evaluated only on audio drama domain; generalizability to other long-form narratives is unknown.
  • Closed-frontier systems saturate at plot-beat F1 [0.78, 0.81] and collapse by -0.20 F1 at horizon h=200, indicating fundamental limitations in long-horizon coherence for current LLMs.

Tags

audio dramalong-horizonbenchmarknarrative understandingstate-space modelLLM evaluationCLAI