4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation Models
TLDR
4DWorldBench is a unified evaluation framework for 3D/4D world generation models, assessing perceptual quality, alignment, physical realism, and consistency.
Reasoning
The paper introduces a comprehensive benchmark for world generation models, covering multiple tasks and evaluation dimensions, which is a strength. However, it focuses solely on generation and does not address interactive or RL-based world models, limiting its scope.
Read-first score
Read-first score 58.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 37.
Field roles
Rank sensitivity
Stability: volatile; rank range: 273.
Keyword Scores
Deep Analysis
Innovations
- Comprehensive evaluation framework for 3D/4D world generation across four key dimensions: Perceptual Quality, Condition-4D Alignment, Physical Realism, and 4D Consistency
- Adaptive conditioning across multiple modalities with mapping to a unified textual space
- Integration of LLM-as-judge, MLLM-as-judge, and traditional network-based methods for unified evaluation
- Extension of traditional evaluation paradigms to include adaptive tool selection for closer agreement with human judgments
Methodology
The benchmark evaluates world generation models on tasks such as Image-to-3D/4D, Video-to-4D, and Text-to-3D/4D. All modality conditions are mapped into a unified textual space, and evaluation is performed using LLM-as-judge, MLLM-as-judge, and traditional network-based methods. An adaptive tool selection mechanism is employed to choose the most appropriate evaluation method for each input.
Key Results
Preliminary human studies demonstrate that the adaptive tool selection achieves closer agreement with subjective human judgments compared to fixed evaluation methods.
Limitations
- Preliminary human studies only, not a full-scale validation
- Potential biases inherent in LLM/MLLM-based judges are not explicitly addressed