Evaluating the World Model Implicit in a Generative Model
TLDR
Recent work suggests that large language models may implicitly learn world models.
Reasoning
Fallback reasoning generated from available title and abstract metadata: Recent work suggests that large language models may implicitly learn world models. How should we assess this possibility? We formalize this question for the case where the underlying reality is governed by a deterministic finite automaton. This includes...
Read-first score
Read-first score 55.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals.
Field roles
Rank sensitivity
Stability: volatile; rank range: 637.
Deep Analysis
Innovations
- Formalizing the evaluation of world models in generative models for deterministic finite automaton domains
- Proposing new evaluation metrics inspired by the Myhill-Nerode theorem from language theory
- Demonstrating that generative models can pass existing diagnostics while having incoherent world models, leading to fragility
Methodology
The authors formalize world model recovery for domains governed by deterministic finite automata. They introduce evaluation metrics based on the Myhill-Nerode theorem and test them on three domains: game playing, logic puzzles, and navigation. They compare performance on existing diagnostics with their new metrics to assess coherence.
Key Results
Generative models perform well on existing diagnostics for world models, but the proposed metrics reveal that their world models are far less coherent than they appear, causing fragility when solving related but subtly different tasks.
Limitations
- Formalization is limited to deterministic finite automaton, which may not cover all real-world world model scenarios
- Evaluation is only conducted on three domains (game playing, logic puzzles, navigation), leaving generalizability unaddressed
- The new metrics assess coherence but may not capture all aspects of world model accuracy or completeness