NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama
TLDR
A benchmark and latent world model for long-horizon audio drama, showing frontier LLMs saturate while N-VSSM maintains high consistency across 200 episodes.
Reasoning
Strengths include a novel benchmark with cross-lingual evaluation and a latent world model that outperforms frontier LLMs on long-arc consistency with lower compute, validated by a human study. Weaknesses are domain specificity to audio drama and lack of generalization evidence to other modalities.
Read-first score
Read-first score 57.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 34.
Field roles
Rank sensitivity
Stability: volatile; rank range: 404.
Keyword Scores
Deep Analysis
Innovations
- NarrativeWorldBench: an open benchmark of nine narrative-structure metrics evaluated across horizons h in {10, 20, 50, 100, 200} with cross-lingual evaluation across four Indic languages (Hindi, Tamil, Telugu, Marathi).
- N-VSSM: a Narrative Variational State-Space Model that maintains a structured 256-dimensional latent world state over more than 200 episodes via a Mamba-2 backbone with an event-conditioned posterior and an 8B decoder.
- A learned Cultural Transfer Function that lifts cross-language fidelity by +0.20 to +0.23 Likert points.
Methodology
The paper benchmarks 21 models across classical, fine-tuned, open-frontier, closed-frontier, and reasoning tiers on a uniform set of structural narrative metrics. It introduces NarrativeWorldBench with nine metrics and cross-lingual evaluation, and proposes N-VSSM, a latent world model using a Mamba-2 backbone and event-conditioned posterior. A within-subjects writer study with 12 professional authors and 240 trials compares N-VSSM against Claude Opus 4.5.
Key Results
N-VSSM achieves plot-beat F1 = 0.84 across all horizons at 4x lower compute than the closed-frontier band. The Cultural Transfer Function improves cross-language fidelity by +0.20 to +0.23 Likert points. In the writer study, N-VSSM is preferred over Claude Opus 4.5 on long-arc consistency 71% of the time and rated +1.3 Likert points higher on controllability.
Limitations
- Benchmark only covers four Indic languages, limiting cross-lingual generality.
- Writer study uses only 12 professional authors, which may not be representative.
- Model's performance is evaluated only on audio drama domain; generalizability to other long-form narratives is unknown.
- Closed-frontier systems saturate at plot-beat F1 [0.78, 0.81] and collapse by -0.20 F1 at horizon h=200, indicating fundamental limitations in long-horizon coherence for current LLMs.