Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models

arXiv 26.3 2026 68 benchmark

TLDR

Omni-WorldBench evaluates world models' interactive response in 4D settings using a prompt suite and agent-based metrics, revealing limitations across 18 models.

Reasoning

The paper addresses a critical gap in evaluating interactive response for world models, offering a comprehensive benchmark with a prompt suite and agent-based metrics. However, it is limited to video-based paradigms and lacks explicit details on real-world dataset composition.

Read-first score

Read-first score 68, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 54.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,benchmark,evaluation,metric

Topical relevance 42%
77.1

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 146.

Keyword Scores

world model
10
interactive world model
10
video world model
9
generative world model
8
world dynamics prediction
8
world simulator
7
model-based reinforcement learning world model
2

Deep Analysis

Innovations

  • Proposes Omni-WorldBench, a comprehensive benchmark specifically designed to evaluate interactive response capabilities of world models in 4D settings.
  • Introduces Omni-WorldSuite, a systematic prompt suite spanning diverse interaction levels and scene types.
  • Introduces Omni-Metrics, an agent-based evaluation framework that quantifies world modeling capabilities by measuring the causal impact of interaction actions on final outcomes and intermediate state evolution trajectories.

Methodology

The benchmark comprises two key components: Omni-WorldSuite (a prompt suite) and Omni-Metrics (an agent-based evaluation framework). It evaluates 18 representative world models across multiple paradigms using these components to measure interactive response capabilities.

Key Results

Extensive evaluations of 18 world models reveal critical limitations of current world models in interactive response, providing actionable insights for future research.

Tags