Diffusion Models for Video Prediction and Infilling
TLDR
RaMViD extends diffusion models to videos with 3D convolutions and masking for state-of-the-art video prediction and infilling.
Reasoning
The paper introduces a novel conditioning technique for video diffusion models, achieving state-of-the-art results on benchmark datasets. However, it focuses narrowly on video prediction and infilling without explicit connection to world models or reinforcement learning, limiting its scope.
Read-first score
Read-first score 32.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 9.
Field roles
Candidate
Rank sensitivity
Stability: volatile; rank range: 49.