Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Yume-1.5: A Text-Controlled Interactive World Generation Model

arXiv 25.12 2025 53 method, system

TLDR

Yume-1.5 generates interactive, explorable worlds from text or image using context compression, streaming acceleration, and text-controlled events.

Reasoning

The paper proposes a novel framework addressing key limitations like large parameters and slow inference, but the abstract lacks explicit evaluation results or benchmark comparisons, making it hard to assess empirical validity.

Read-first score

Read-first score 53, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 40.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Topical relevance 42%
57.1

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 25%
50

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=result

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=code

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 408.

Keyword Scores

interactive world model
9
generative world model
8
video world model
7
world model
6
world simulator
5
world dynamics prediction
4
model-based reinforcement learning world model
1

Deep Analysis

Innovations

  • Long-video generation framework integrating unified context compression with linear attention
  • Real-time streaming acceleration strategy powered by bidirectional attention distillation and enhanced text embedding scheme
  • Text-controlled method for generating world events

Methodology

The proposed framework generates realistic, interactive, and continuous worlds from a single image or text prompt, enabling keyboard-based exploration. It comprises three core components: a long-video generation framework with unified context compression and linear attention, a real-time streaming acceleration strategy using bidirectional attention distillation and enhanced text embeddings, and a text-controlled method for generating world events.

Key Results

The model generates realistic, interactive, and continuous worlds from a single image or text prompt, enabling keyboard-based exploration. No quantitative results are reported in the abstract.

Tags