用于多步视觉推理的分层去噪
TLDR
提出HDR,一种用于视频生成中多步视觉推理的分层去噪框架,实现了推理一致性和低延迟。
评分理由
Strengths: novel hierarchical latent structure enabling coarse-to-fine reasoning, a new benchmark with six diverse tasks, and significant performance gains over baselines. Weaknesses: limited to synthetic visual reasoning tasks, no explicit connection to real-world video data or scalability analysis.
Read-first 评分解释
综合优先阅读分 21.4,由主题、引用、图谱、方法、可复现性和近期性等信号加权得到。 原始总分保留为 0。
研究版图角色
前沿论文
排序敏感性
稳定性:volatile;排名波动范围:37。