Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Reference-Free Assessment of Physical Consistency in World Model-based Video Generation

arXiv 2026 58.4 benchmark

TLDR

Introduces reference-free metrics using DROID-SLAM and SEA-RAFT to evaluate physical consistency in world model-based video generation, improving task success rates by 8%.

Reasoning

Strengths: novel reference-free evaluation method that quantifies physical inconsistencies and improves task success rates. Weaknesses: limited to specific SLAM and optical flow tools, and generalizability to diverse video generation tasks is unclear.

Read-first score

Read-first score 58.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 44.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
94.6

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.94574276

Topical relevance 29%
62.9

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 18%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=evaluation,metric

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=artifact

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridge

Rank sensitivity

Stability: volatile; rank range: 449.

Keyword Scores

world model
9
video world model
9
generative world model
8
world simulator
6
world dynamics prediction
5
interactive world model
4
model-based reinforcement learning world model
3

Deep Analysis

Innovations

  • Reference-free measures for evaluating physical consistency of generated videos
  • Combining relative and absolute approaches to assess fidelity
  • Use of DROID-SLAM and SEA-RAFT to quantify physical inconsistencies
  • Spatio-temporal localization of physical artifacts via absolute assessment

Methodology

The paper introduces reference-free measures combining relative and absolute approaches to evaluate physical consistency in world model-based video generation. The relative consistency assessment uses DROID-SLAM and SEA-RAFT to quantify physical inconsistencies, motivated by WorldScore, while the absolute assessment enables spatio-temporal localization of artifacts. The method is evaluated by filtering videos based on relative consistency and measuring task success rate improvements.

Key Results

Videos filtered using the relative consistency assessment show an improvement in task success rates of over 8%, effectively narrowing the simulation-to-reality gap. The absolute assessment provides spatio-temporal localization, visualizing when and where physical artifacts occur.

Tags

physical consistencyvideo generationworld modelsevaluation metricssimulation-to-reality gapAILGRO