WorldOdysseyBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
TLDR
Introduces WorldOdysseyBench, a benchmark for long-horizon stability of interactive world models across action, vision, physics, and memory dimensions.
Reasoning
Strengths include a comprehensive multi-dimensional benchmark with novel metrics and diverse real-world scenes; weaknesses are that the abstract does not detail specific model failures or limitations beyond moderate scores.
Read-first score
Read-first score 39.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 39.
Field roles
Rank sensitivity
Stability: volatile; rank range: 220.
Keyword Scores
Deep Analysis
Innovations
- Per-frame action metric that bypasses cross-model semantic scale disparity and exposes failures hidden by trajectory-level evaluation
- Segment-based drift metric that captures non-monotonic mid-sequence collapse missed by start-vs-end comparisons
- Controllability-gated evaluation over mechanics, optics, and 3D consistency, scoring plausibility under faithful action execution
- Action-decoupled memory protocol evaluating scene memory via transition-localized 3D point-cloud reconstruction and subject memory via tracking-plus-VLM reasoning
Methodology
WorldOdysseyBench is a benchmark comprising 600+ test cases across Nature, Urban, and Indoor scenes in first/third-person views with WASD continuous interaction lasting 10-60 seconds. It evaluates interactive world models on four dimensions—Action, Vision, Physics, and Memory—using tailored metrics and protocols, and tests over 10 open- and closed-source models.
Key Results
No evaluated model reliably satisfies all four dimensions; even the best-performing model achieves only moderate scores.