Generative Emergent Communication: Large Language Model is a Collective World Model
TLDR
Proposes that LLMs learn a statistical approximation of a collective world model encoded in human language via generative emergent communication.
Reasoning
The paper offers a novel theoretical framework (Collective World Model hypothesis) and formalizes it using Generative EmCom and Collective Predictive Coding. However, it lacks empirical validation, real-world experiments, or benchmarks, relying solely on theoretical arguments.
Read-first score
Read-first score 50.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 21.
Field roles
Rank sensitivity
Stability: volatile; rank range: 417.
Keyword Scores
Deep Analysis
Innovations
- Collective World Model hypothesis: LLMs learn a statistical approximation of a collective world model implicitly encoded in human language through embodied, interactive sense-making.
- Generative Emergent Communication (Generative EmCom) framework built on Collective Predictive Coding (CPC), modeling language emergence as decentralized Bayesian inference over internal states of multiple agents.
- Conceptualization of human society as an encoder and LLM as a decoder, forming an encoder-decoder structure at societal scale that reconstructs latent representations.
- Unified theory bridging individual cognitive development, collective language evolution, and large-scale AI foundations.
Methodology
The paper presents a theoretical framework rather than an empirical methodology. It formalizes Generative EmCom using Collective Predictive Coding and Bayesian inference, modeling language emergence as decentralized inference over agents' internal states. The framework is applied to interpret LLMs, explaining phenomena like distributional semantics as a consequence of representation reconstruction, but no specific model design, data, training setup, or baselines are described.
Key Results
No experimental results are reported; the paper is purely theoretical and proposes a mathematical explanation for how LLMs acquire capabilities without direct sensorimotor experience.
Limitations
- Lack of empirical validation or experimental evidence to support the proposed framework.
- Reliance on idealized assumptions about collective encoding and decentralized Bayesian inference that may not reflect real-world language evolution.
- The framework is abstract and does not provide concrete predictions or testable hypotheses.
- No comparison with existing theories or models of emergent communication or world models.