Video-GPT via Next Clip Diffusion
TLDR
Proposes Video-GPT using next clip diffusion for video as language, achieving SOTA on video prediction for world modeling.
Reasoning
Strengths include a novel paradigm and strong benchmark results on Physics-IQ. Weaknesses are limited methodological details and unclear distinction from video prediction alone.
Read-first score
Read-first score 44, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 29.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 370.