Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
TLDR
Divot introduces a diffusion-powered video tokenizer that unifies video comprehension and generation in LLMs, achieving competitive benchmark performance.
Reasoning
The paper presents a novel tokenizer leveraging diffusion for self-supervised video representation learning and a diffusion-based de-tokenizer, with strong empirical results on video benchmarks. However, the abstract does not explicitly connect the method to world models or dynamics prediction, limiting relevance to those keywords.
Read-first score
Read-first score 35.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 9.
Field roles
Candidate
Rank sensitivity
Stability: volatile; rank range: 79.