WorldMark: A Unified Benchmark Suite for Interactive Video World Models
TLDR
WorldMark is the first unified benchmark for interactive video world models, enabling fair comparison via standardized scenes, actions, and evaluation metrics.
Reasoning
The paper addresses a clear gap in the field by providing a common evaluation framework for interactive video world models, with strong methodological contributions like a unified action-mapping layer and hierarchical test suite. However, the abstract does not detail specific experimental results or limitations, and the claim of being 'first' may need verification against prior work.
Read-first score
Read-first score 61.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 51.
Field roles
Rank sensitivity
Stability: volatile; rank range: 375.
Keyword Scores
Deep Analysis
Innovations
- Unified action-mapping layer that translates a shared WASD-style action vocabulary into each model's native control format, enabling apples-to-apples comparison across six major models on identical scenes and trajectories.
- Hierarchical test suite of 500 evaluation cases covering first- and third-person viewpoints, photorealistic and stylized scenes, and three difficulty tiers spanning 20-60 seconds.
- Modular evaluation toolkit for Visual Quality, Control Alignment, and World Consistency, designed for reuse of standardized inputs while allowing plug-in of custom metrics.
Methodology
WorldMark provides a common playing field for interactive Image-to-Video world models by standardizing scenes, action sequences, and control interfaces. It introduces a unified action-mapping layer to translate a shared WASD-style action vocabulary into each model's native control format, enabling comparison across six models. The benchmark includes a hierarchical test suite of 500 evaluation cases with varied viewpoints, styles, and difficulty tiers, and a modular evaluation toolkit measuring Visual Quality, Control Alignment, and World Consistency.
Key Results
The abstract does not report specific experimental results; it describes the benchmark design, the planned release of all data, evaluation code, and model outputs, and the launch of World Model Arena (warena.ai) for online side-by-side model comparisons and live leaderboard.