PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
TLDR
Introduces PlayWorld, a benchmark with 171 scenarios and agent players to evaluate video world models on long-horizon interactive objectives across geometry, interaction, and evolution dimensions.
Reasoning
The paper addresses a clear evaluation gap by using agent players for cross-model comparison and proposes a multi-dimensional benchmark tested on nine world models. However, the abstract is truncated and lacks detailed quantitative results or explicit limitations, making full assessment difficult.
Read-first score
Read-first score 47.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 47.
Field roles
Rank sensitivity
Stability: volatile; rank range: 581.
Keyword Scores
Deep Analysis
Innovations
- Using multi-modal Agent Players to interact with world models for long-horizon objective evaluation, overcoming fixed action-conditioned comparison issues
- Introducing PlayWorld benchmark with 171 scenarios each having a specified objective
- Defining four core evaluation dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution
- Incorporating basic ability metrics for video quality and controllability
Methodology
The benchmark employs multi-modal Agent Players that interact with world models to pursue long-horizon objectives. It provides 171 scenarios, each with a specified objective. Models are evaluated on four core dimensions (geometry consistency, interaction fidelity, out-of-sight evolution, insight evolution) and basic video quality/controllability metrics, tested on nine state-of-the-art world models.
Key Results
Current world models are unreliable on long-horizon interactive objectives, particularly struggling with spatial consistency and persistent state evolution.