VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation
TLDR
VideoGPA distills geometry priors via self-supervised preference alignment to improve 3D consistency, temporal stability, and motion coherence in video diffusion models.
Reasoning
The paper presents a novel data-efficient framework using geometry foundation models and DPO to enforce 3D consistency in video generation. Strengths include a self-supervised pipeline avoiding human annotations and strong empirical gains; weaknesses are that the abstract provides limited detail on baseline comparisons and evaluation metrics.
Read-first score
Read-first score 23.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 5.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 70.