Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model
TLDR
Introduces an instruction-following interactive game world model using natural language, keyboard, or mouse control, with a benchmark for evaluation.
Reasoning
The paper presents a novel paradigm for interactive game world modeling with flexible instruction-driven control, supported by an automated dataset creation process and a dedicated benchmark. Strengths include the integration of diverse input modalities and causal grounding, while weaknesses may involve scalability and reliance on a large MoE model.
Read-first score
Read-first score 69.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 51.
Field roles
Rank sensitivity
Stability: volatile; rank range: 39.
Keyword Scores
Deep Analysis
Innovations
- Instruction-driven interaction paradigm for generative game world modeling, replacing fixed keyboard inputs with natural language, keyboard, or mouse signals
- Automated process to transform large-scale unstructured text-video pairs into causally aligned interactive datasets
- Text-driven interaction injection mechanism for fine-grained control over camera motion, character behavior, and environment dynamics
- Interaction-focused benchmark InterBench for comprehensive evaluation of interaction performance
Methodology
The model is built upon a 14B image-to-video Mixture-of-Experts (MoE) foundation model, incorporating a text-driven interaction injection mechanism for fine-grained control. An automated process transforms large-scale unstructured text-video pairs into causally aligned interactive datasets. Evaluation is performed using a newly introduced interaction-focused benchmark, InterBench.
Key Results
The model generates temporally coherent and causally grounded interactive game videos that faithfully respond to diverse free-form user instructions such as 'open the door', 'draw a torch', or 'trigger an explosion'.