VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?
TLDR
Proposes VibeWorlding, a benchmark and RL training framework for multimodal agents to construct interactive 3D open worlds from user queries; current MLLMs underperform.
Reasoning
The paper introduces a substantial benchmark and training environment for multimodal 3D world construction, with empirical evaluation of frontier models. However, its core contribution is agent benchmarking and RL post-training rather than world modeling or dynamics prediction, so world-model-related keywords are only weakly relevant.
Read-first score
Read-first score 24.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 11.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 46.