Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

WorldMark: A Unified Benchmark Suite for Interactive Video World Models

arXiv 2026 61.1 benchmark

TLDR

WorldMark is the first unified benchmark for interactive video world models, enabling fair comparison via standardized scenes, actions, and evaluation metrics.

Reasoning

The paper addresses a clear gap in the field by providing a common evaluation framework for interactive video world models, with strong methodological contributions like a unified action-mapping layer and hierarchical test suite. However, the abstract does not detail specific experimental results or limitations, and the claim of being 'first' may need verification against prior work.

Read-first score

Read-first score 61.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 51.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 18%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,experiment,metric,result

Topical relevance 29%
72.9

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Citation impact 18%
63.6

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.63648145

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 375.

Keyword Scores

world model
10
interactive world model
10
video world model
9
generative world model
8
world simulator
6
world dynamics prediction
5
model-based reinforcement learning world model
3

Deep Analysis

Innovations

  • Unified action-mapping layer that translates a shared WASD-style action vocabulary into each model's native control format, enabling apples-to-apples comparison across six major models on identical scenes and trajectories.
  • Hierarchical test suite of 500 evaluation cases covering first- and third-person viewpoints, photorealistic and stylized scenes, and three difficulty tiers spanning 20-60 seconds.
  • Modular evaluation toolkit for Visual Quality, Control Alignment, and World Consistency, designed for reuse of standardized inputs while allowing plug-in of custom metrics.

Methodology

WorldMark provides a common playing field for interactive Image-to-Video world models by standardizing scenes, action sequences, and control interfaces. It introduces a unified action-mapping layer to translate a shared WASD-style action vocabulary into each model's native control format, enabling comparison across six models. The benchmark includes a hierarchical test suite of 500 evaluation cases with varied viewpoints, styles, and difficulty tiers, and a modular evaluation toolkit measuring Visual Quality, Control Alignment, and World Consistency.

Key Results

The abstract does not report specific experimental results; it describes the benchmark design, the planned release of all data, evaluation code, and model outputs, and the launch of World Model Arena (warena.ai) for online side-by-side model comparisons and live leaderboard.

Tags

benchmarkworld modelsinteractive video generationevaluationaction mappingcomputer visionCV