VideoPhy: Evaluating Physical Commonsense for Video Generation
TLDR
A benchmark to evaluate if text-to-video models follow physical commonsense for real-world activities, finding current models severely lacking.
Reasoning
The paper introduces a novel benchmark with human evaluation, highlighting a clear limitation in video generation models. Strengths include a well-defined evaluation setup and actionable results. Weaknesses are the reliance on human evaluation and limited scope of prompts, but the work is sound.
Read-first score
Read-first score 38.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 25.
Field roles
Candidate
Rank sensitivity
Stability: volatile; rank range: 188.