HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
TLDR
HY-World 2.0 is a multi-modal world model that generates, reconstructs, and simulates 3D worlds from text, images, and videos, achieving state-of-the-art performance on benchmarks.
Reasoning
The paper presents a comprehensive framework with clear methodological stages and innovations, and it validates performance on benchmarks, indicating strong empirical grounding. However, the abstract lacks details on limitations and specific benchmark results, and the connection to world dynamics prediction or RL is not evident.
Read-first score
Read-first score 43.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 45.
Field roles
Rank sensitivity
Stability: volatile; rank range: 350.
Keyword Scores
Deep Analysis
Innovations
- HY-Pano 2.0 for enhanced panorama fidelity in world generation
- WorldNav for trajectory planning enabling 3D scene understanding and navigation
- WorldStereo 2.0 with consistent memory for keyframe-based view generation
- WorldMirror 2.0 with refined architecture and learning strategy for 3D reconstruction from multi-view images or videos
- WorldLens, a high-performance 3DGS rendering platform with engine-agnostic architecture, automatic IBL lighting, collision detection, and training-rendering co-design
Methodology
HY-World 2.0 is a multi-modal world model that accepts text, single-view images, multi-view images, or videos, and outputs 3D Gaussian Splatting scenes. For generative inputs (text or single image), it uses a four-stage pipeline: panorama generation (HY-Pano 2.0), trajectory planning (WorldNav), world expansion (WorldStereo 2.0), and world composition (WorldMirror 2.0). WorldLens provides interactive rendering with character support.
Key Results
HY-World 2.0 achieves state-of-the-art performance on multiple benchmarks among open-source approaches and delivers results comparable to the closed-source model Marble.