GODIVA:从自然描述生成开放域视频
TLDR
GODIVA是开放域文本生成视频预训练模型,采用3D稀疏注意力自回归生成,在Howto100M上预训练,并提出RM评估指标。
评分理由
The paper presents a text-to-video generation model with large-scale pretraining and zero-shot evaluation, which are strengths. However, it does not address world models or interactive dynamics, and the abstract lacks details on limitations and baseline comparisons.
Read-first 评分解释
综合优先阅读分 29.2,由主题、引用、图谱、方法、可复现性和近期性等信号加权得到。 原始总分保留为 1。
研究版图角色
候选论文
排序敏感性
稳定性:volatile;排名波动范围:54。