OSCBench: Benchmarking Object State Change in Text-to-Video Generation
TLDR
Introduces OSCBench, a benchmark evaluating object state changes in text-to-video generation; models struggle with accurate state transformations.
Reasoning
Strengths: addresses an underexplored aspect of action understanding, uses systematic categorization and dual evaluation (human+MLLM). Weaknesses: limited to instructional cooking data and only six models, potentially narrow domain; no discussion of broader generalization beyond cooking.
Read-first score
Read-first score 22.9, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 7.
Field roles
Frontier
Rank sensitivity
Stability: volatile; rank range: 29.