Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models

arXiv 25.11 2025 58.8 benchmark

TLDR

4DWorldBench is a unified evaluation framework for 3D/4D world generation models, assessing perceptual quality, alignment, physical realism, and consistency.

Reasoning

The paper introduces a comprehensive benchmark for world generation models, covering multiple tasks and evaluation dimensions, which is a strength. However, it focuses solely on generation and does not address interactive or RL-based world models, limiting its scope.

Read-first score

Read-first score 58.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 37.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Methodology quality 25%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,evaluation,validation

Topical relevance 42%
52.9

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=github

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 273.

Keyword Scores

world model
9
generative world model
8
video world model
7
world dynamics prediction
6
world simulator
4
interactive world model
2
model-based reinforcement learning world model
1

Deep Analysis

Innovations

  • Comprehensive evaluation framework for 3D/4D world generation across four key dimensions: Perceptual Quality, Condition-4D Alignment, Physical Realism, and 4D Consistency
  • Adaptive conditioning across multiple modalities with mapping to a unified textual space
  • Integration of LLM-as-judge, MLLM-as-judge, and traditional network-based methods for unified evaluation
  • Extension of traditional evaluation paradigms to include adaptive tool selection for closer agreement with human judgments

Methodology

The benchmark evaluates world generation models on tasks such as Image-to-3D/4D, Video-to-4D, and Text-to-3D/4D. All modality conditions are mapped into a unified textual space, and evaluation is performed using LLM-as-judge, MLLM-as-judge, and traditional network-based methods. An adaptive tool selection mechanism is employed to choose the most appropriate evaluation method for each input.

Key Results

Preliminary human studies demonstrate that the adaptive tool selection achieves closer agreement with subjective human judgments compared to fixed evaluation methods.

Limitations

  • Preliminary human studies only, not a full-scale validation
  • Potential biases inherent in LLM/MLLM-based judges are not explicitly addressed

Tags