Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

LoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)

arXiv 2026 55.5 benchmark

TLDR

The LoViF 2026 PhyScore challenge benchmarks holistic quality assessment of world-model-generated videos, evaluating physical realism, temporal consistency, and anomaly localization.

Reasoning

Strengths: addresses the gap in evaluating physical plausibility beyond perceptual quality, with a comprehensive benchmark of 1,554 videos across multiple tracks and human annotations. Weaknesses: focuses solely on evaluation metrics rather than proposing new world models or generation methods; abstract lacks results or analysis of submitted solutions.

Read-first score

Read-first score 55.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 34.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 18%
90

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,dataset,evaluation,metric,result

Citation impact 18%
72.2

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.721973

Topical relevance 29%
48.6

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FrontierBridgeMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 379.

Keyword Scores

world model
8
generative world model
7
world dynamics prediction
7
video world model
6
world simulator
4
interactive world model
1
model-based reinforcement learning world model
1

Deep Analysis

Innovations

  • First challenge on holistic quality assessment for 4D world models (PhyScore), addressing the gap in evaluating physical plausibility beyond perceptual quality.
  • Joint prediction of four quality dimensions (Video Quality, Physical Realism, Condition-Video Alignment, Temporal Consistency) combined with physical anomaly timestamp localization.
  • Benchmark dataset of 1,554 videos from seven world generative models across three tracks (text-2D, image-to-4D, video-to-4D) and 26 categories covering physics-relevant scenarios (dynamics, optics, thermodynamics).
  • Composite evaluation protocol combining TimeStamp_IOU for anomaly localization and SRCC/PLCC for score prediction.

Methodology

Participants are required to build a metric that jointly predicts four dimensions (Video Quality, Physical Realism, Condition-Video Alignment, Temporal Consistency) and localizes physical anomaly timestamps. The benchmark dataset contains 1,554 videos generated by seven world generative models across three tracks and 26 categories. Labels are produced through trained human annotation with an automated quality-control pass. Evaluation uses a composite protocol combining TimeStamp_IOU for anomaly localization and SRCC/PLCC for score prediction.

Key Results

The abstract does not report specific quantitative results; it summarizes the challenge design and provides method-level insights from submitted solutions.

Tags

video quality assessmentphysical realismtemporal consistencyworld modelbenchmarkCV