A Unified Definition of Hallucination, Or: It's the World Model, Stupid
TLDR
Despite numerous attempts at mitigation since the inception of language models, hallucinations remain a persistent problem even in today's frontier LLMs.
Reasoning
Fallback reasoning generated from available title and abstract metadata: Despite numerous attempts at mitigation since the inception of language models, hallucinations remain a persistent problem even in today's frontier LLMs. Why is this? We review existing definitions of hallucination and fold them into a single, unified definition...
Read-first score
Read-first score 40.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals.
Field roles
Rank sensitivity
Stability: volatile; rank range: 215.
Deep Analysis
Innovations
- Unified definition of hallucination as inaccurate internal world modeling observable to the user
- Framework that subsumes prior definitions by varying reference world model and conflict policy
- Distinction between true hallucinations and planning or reward errors
- Common language for comparison across benchmarks and discussion of mitigation strategies
- Connection to HalluWorld benchmark for stress-testing model hallucinations
Methodology
The paper reviews existing definitions of hallucination and synthesizes them into a unified definition based on inaccurate world modeling. It proposes a framework that varies the reference world model and conflict policy to subsume prior definitions. The work also connects this framework to the HalluWorld benchmark, which instantiates fully specified reference world models for stress-testing.
Key Results
No experimental results are presented; the paper is a theoretical/definitional work that introduces a unified definition and connects it to the HalluWorld benchmark.