Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Trimming the Long-Tail of Visual World Modeling Evaluation

arXiv 2026 60.2 benchmark

TLDR

Introduces Tailor-Bench to evaluate visual world models on long-tail physical interactions, revealing limited generalization beyond common scenarios.

Reasoning

The paper presents a well-structured benchmark with three scenario modes and two generation settings, effectively demonstrating a long-tail gap in physical world modeling. However, it focuses solely on evaluation without proposing new models or addressing interactive or reinforcement learning contexts.

Read-first score

Read-first score 60.2, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 39.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
94.8

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.94788889

Methodology quality 18%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=analysis,benchmark,evaluation,experiment,result

Topical relevance 29%
55.7

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 418.

Keyword Scores

world model
9
generative world model
8
video world model
7
world dynamics prediction
6
world simulator
5
interactive world model
3
model-based reinforcement learning world model
1

Deep Analysis

Innovations

  • Introduction of Tailor-Bench, a benchmark for evaluating visual world models on irregular physical interactions
  • Three scenario modes (Regular, Unconventional, Impossible) to progressively challenge model reasoning
  • Two complementary evaluation settings (predictive generation and descriptive generation) under a unified protocol

Methodology

Tailor-Bench is designed with three scenario modes: Regular (common tool-task pairs), Unconventional (attribute-compatible substitutes), and Impossible (attribute-violating tools). Two settings are used: predictive generation (infer outcomes without guidance) and descriptive generation (specify target outcome). The benchmark evaluates image and video generation models on physical interaction simulation.

Key Results

Performance degrades from Regular to Unconventional to Impossible scenarios, revealing a long-tail gap in physical world modeling. Image models fail to realize correct state changes, while video models suffer from temporal inconsistencies.

Tags

computer visionworld modelsbenchmarkphysical reasoninglong-tail distributionevaluationCV