DeRA: Decoupled Representation Alignment for Video Tokenization
TLDR
DeRA is a 1D video tokenizer that decouples spatial and temporal representation learning, aligning with vision foundation models to improve video generation efficiency and performance.
Reasoning
The paper presents a novel video tokenizer with a decoupled architecture and a gradient conflict resolution module, showing strong empirical results on video generation benchmarks. However, the abstract does not position the work as a world model, so relevance to world model keywords is limited to indirect connections via video generation.
Read-first score
Read-first score 29.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 4.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 118.