DiLA: Disentangled Latent Action World Models
TLDR
DiLA introduces a disentangled latent action world model that resolves the trade-off between action abstraction and video generation fidelity via content-structure disentanglement.
Reasoning
The paper presents a novel approach to latent action models by co-evolving disentanglement and latent action learning, achieving high-quality video generation and interpretable action spaces. Strengths include clear problem formulation and strong empirical results across multiple tasks; weaknesses are the lack of explicit real-world dataset names and potential limitations in scalability or domain generality.
Read-first score
Read-first score 55.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 51.
Field roles
Rank sensitivity
Stability: volatile; rank range: 432.
Keyword Scores
Deep Analysis
Innovations
- Disentangled latent action world model (DiLA) that resolves the trade-off between action abstraction and generation fidelity via content-structure disentanglement
- Co-evolving disentanglement and latent action learning, where the predictive bottleneck drives separation of spatial layouts (structure) and visual details (content)
- Continuous, semantically structured latent action space that maintains high generative quality
Methodology
DiLA employs a content-structure disentanglement framework within a latent action world model. The predictive bottleneck inherent in latent action learning forces the model to distill spatial layouts into a structure pathway while offloading visual details to a separate content pathway, enabling simultaneous high-level action abstraction and high-fidelity generation.
Key Results
DiLA achieves superior results in video generation quality, action transfer, visual planning, and manifold interpretability compared to existing methods.