VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
TLDR
VideoREPA distills physics understanding from video foundation models into text-to-video diffusion models via token relation alignment, improving physical plausibility of generated videos.
Reasoning
The paper proposes a novel token relation distillation loss to inject physics knowledge into T2V models and reports benchmark improvements over CogVideoX. Strengths include a clear problem framing and a new alignment method, but the abstract lacks quantitative details and broader comparisons, and the connection to world models is not explicitly established.
Read-first score
Read-first score 34.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 4.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 203.