Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

DiLA: Disentangled Latent Action World Models

arXiv 2026 55.4 method

TLDR

DiLA introduces a disentangled latent action world model that resolves the trade-off between action abstraction and video generation fidelity via content-structure disentanglement.

Reasoning

The paper presents a novel approach to latent action models by co-evolving disentanglement and latent action learning, achieving high-quality video generation and interpretable action spaces. Strengths include clear problem formulation and strong empirical results across multiple tasks; weaknesses are the lack of explicit real-world dataset names and potential limitations in scalability or domain generality.

Read-first score

Read-first score 55.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 51.

Recency 6%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Citation impact 18%
79.2

Uses OpenAlex-shaped citation metadata as a bibliometric attention signal, separate from paper quality. citation_normalized_percentile=0.79203302

Topical relevance 29%
72.9

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 18%
50

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=result

Reproducibility 18%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Citation velocity 12%
0

Citation velocity estimates citations per publication-year to reduce old-paper bias. velocity=0.00

Field roles

FoundationFrontierBridge

Rank sensitivity

Stability: volatile; rank range: 432.

Keyword Scores

world model
10
video world model
9
generative world model
8
world dynamics prediction
8
interactive world model
6
world simulator
5
model-based reinforcement learning world model
5

Deep Analysis

Innovations

  • Disentangled latent action world model (DiLA) that resolves the trade-off between action abstraction and generation fidelity via content-structure disentanglement
  • Co-evolving disentanglement and latent action learning, where the predictive bottleneck drives separation of spatial layouts (structure) and visual details (content)
  • Continuous, semantically structured latent action space that maintains high generative quality

Methodology

DiLA employs a content-structure disentanglement framework within a latent action world model. The predictive bottleneck inherent in latent action learning forces the model to distill spatial layouts into a structure pathway while offloading visual details to a separate content pathway, enabling simultaneous high-level action abstraction and high-fidelity generation.

Key Results

DiLA achieves superior results in video generation quality, action transfer, visual planning, and manifold interpretability compared to existing methods.

Tags

latent action modelsworld modelsdisentangled representationvideo predictionunsupervised learningCVAIRO