MCVD: Masked Conditional Video Diffusion for Prediction, Generation, and Interpolation
TLDR
MCVD uses masked conditional video diffusion to perform video prediction, generation, and interpolation with a single model, achieving SOTA results.
Reasoning
The paper introduces a novel masking technique that enables a single diffusion model to handle multiple video synthesis tasks, with strong empirical results on standard benchmarks. However, it does not address long-term temporal consistency or explicitly model world dynamics, limiting its scope as a world model.
Read-first score
Read-first score 30.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 3.
Field roles
Candidate
Rank sensitivity
Stability: volatile; rank range: 27.