Benchmarking World-Model Learning
Dataset Analysis
Proposes WorldTest protocol and AutumnBench benchmark to evaluate world models on multiple environment-level queries, showing humans outperform frontier models.
Provenance
Collected from papers.
Derived from paper: Benchmarking World-Model Learning