Can Language Models Serve as Text-Based World Simulators?
TLDR
LLMs are tested as text-based world simulators using a new benchmark; GPT-4 proves unreliable, highlighting limitations.
Reasoning
The paper introduces a novel benchmark (ByteSized32-State-Prediction) to directly quantify LLM performance as world simulators, which is a strength. However, it only tests GPT-4 and focuses on text-based environments, limiting generalizability. The empirical evaluation provides concrete evidence for its claims.
Read-first score
Read-first score 67.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 40.
Field roles
Rank sensitivity
Stability: volatile; rank range: 252.
Keyword Scores
Deep Analysis
Innovations
- Introduction of ByteSized32-State-Prediction, a new benchmark for quantifying LLMs as text-based world simulators
- First direct quantification of how well LLMs can serve as text-based world simulators
- Insights into GPT-4's capabilities and weaknesses as a world simulator
Methodology
The authors constructed a new benchmark, ByteSized32-State-Prediction, consisting of a dataset of text game state transitions and accompanying game tasks. They used this benchmark to evaluate GPT-4's ability to predict state changes, directly quantifying its performance as a text-based world simulator.
Key Results
GPT-4, despite impressive performance, is still an unreliable world simulator without further innovations.
Limitations
- GPT-4 is unreliable as a world simulator, indicating current LLMs are insufficient without further innovations