Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Evaluating the World Model Implicit in a Generative Model

arXiv 24.6 2024 55.5 method, theory

TLDR

Recent work suggests that large language models may implicitly learn world models.

Reasoning

Fallback reasoning generated from available title and abstract metadata: Recent work suggests that large language models may implicitly learn world models. How should we assess this possibility? We formalize this question for the case where the underlying reality is governed by a deterministic finite automaton. This includes...

Read-first score

Read-first score 55.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals.

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=evaluation,metric,result

Recency 8%
75.1

Uses a gentle age decay so recent papers surface without erasing older foundations. 2024

Reproducibility 25%
73

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=True; dataset=False; markers=github

Topical relevance 42%
26.5

Matches configured research keywords against title, abstract, tags, and analysis text. matched=5

Field roles

Methodology anchorReproducibility anchor

Rank sensitivity

Stability: volatile; rank range: 637.

Deep Analysis

Innovations

  • Formalizing the evaluation of world models in generative models for deterministic finite automaton domains
  • Proposing new evaluation metrics inspired by the Myhill-Nerode theorem from language theory
  • Demonstrating that generative models can pass existing diagnostics while having incoherent world models, leading to fragility

Methodology

The authors formalize world model recovery for domains governed by deterministic finite automata. They introduce evaluation metrics based on the Myhill-Nerode theorem and test them on three domains: game playing, logic puzzles, and navigation. They compare performance on existing diagnostics with their new metrics to assess coherence.

Key Results

Generative models perform well on existing diagnostics for world models, but the proposed metrics reveal that their world models are far less coherent than they appear, causing fragility when solving related but subtly different tasks.

Limitations

  • Formalization is limited to deterministic finite automaton, which may not cover all real-world world model scenarios
  • Evaluation is only conducted on three domains (game playing, logic puzzles, navigation), leaving generalizability unaddressed
  • The new metrics assess coherence but may not capture all aspects of world model accuracy or completeness

Tags