Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

EchoWM: Open and Enterable Omnimodal World Models

arXiv 2026 48.1 method

TLDR

EchoWM is an omnimodal world model for enterable generative media, jointly generating 720p video, audio, music, and speech while following continuous 6-DoF navigation in first- and third-person scenes.

Reasoning

The paper presents a strong integration of multimodal generation and interactive trajectory control, with evaluations on public world-model benchmarks. However, the abstract lacks specific quantitative results and a clear discussion of limitations, making it hard to fully assess robustness and generalizability.

Read-first score

Read-first score 48.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 51.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Methodology quality 18%
80

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=benchmark,dataset,evaluation,metric

Topical relevance 29%
72.9

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Reproducibility 18%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=dataset

Citation impact 18%
0

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. cited_by_count=0

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FrontierMethodology anchor

Rank sensitivity

Stability: volatile; rank range: 523.

Keyword Scores

world model
10
generative world model
9
interactive world model
9
world simulator
8
video world model
8
world dynamics prediction
7
model-based reinforcement learning world model
0

Deep Analysis

Innovations

  • Omnimodal world model jointly generating 720p video, environmental sound, music, and speech
  • Camera intent-based interaction: first-person observer motion, third-person learned camera-character dynamics without view-specific controllers
  • Mapping discrete commands and continuous poses to a shared metric-scale relative 6-DoF trajectory with dataset-level calibration for consistent motion magnitude
  • Progressive training followed by autoregressive post-training for long-horizon generation

Methodology

The model organizes interaction around camera intent, mapping discrete/continuous inputs to a shared metric-scale 6-DoF trajectory with dataset-level calibration. A complementary data engine is constructed, and progressive training is used, followed by autoregressive post-training for long-horizon generation.

Key Results

EchoWM achieves strong trajectory following and high visual quality on public benchmarks, supports both first- and third-person interaction, and maintains synchronized environmental sound and speech over long-horizon generation.

Tags