AGI Maze as a Benchmark Framework for World-Modeling Agents
TLDR
Introduces AGI Maze, a benchmark for evaluating world-modeling in LLMs via grid-based mazes requiring memory and hidden state representation.
Reasoning
The paper presents a novel benchmark framework for testing world-modeling capabilities, with clear methodology and initial results showing LLM limitations. However, it lacks real-world experiments and the baseline agent evaluation is preliminary, limiting generalizability.
Read-first score
Read-first score 30.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 26.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 112.