Text2World: Benchmarking Large Language Models for Symbolic World Model Generation
TLDR
Introduces Text2World, a PDDL-based benchmark for evaluating LLMs on symbolic world model generation, revealing limited capabilities despite RL-trained reasoning models.
Reasoning
The paper addresses a clear gap with a novel benchmark using execution-based metrics across diverse domains, but the focus is on symbolic models rather than real-world or interactive environments, limiting direct applicability. Strengths include robust evaluation and insights into LLM limitations; weaknesses include lack of real-world validation and narrow domain scope.
Read-first score
Read-first score 57, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 34.
Field roles
Rank sensitivity
Stability: volatile; rank range: 308.
Keyword Scores
Deep Analysis
Innovations
- Introduction of Text2World, a novel benchmark based on PDDL with hundreds of diverse domains
- Multi-criteria, execution-based metrics for robust evaluation of world model generation
- Systematic benchmarking of current LLMs, including reasoning models trained with large-scale reinforcement learning
- Analysis of promising strategies such as test-time scaling and agent training to enhance world modeling
Methodology
The paper introduces Text2World, a benchmark built on the Planning Domain Definition Language (PDDL), comprising hundreds of diverse domains. It employs multi-criteria, execution-based metrics to evaluate LLMs' ability to generate symbolic world models from textual descriptions. Current LLMs are benchmarked, and strategies like test-time scaling and agent training are examined.
Key Results
Reasoning models trained with large-scale reinforcement learning outperform other LLMs on Text2World, but even the best-performing model exhibits limited world modeling capabilities.