PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
TLDR
Introduces PAWBench and PAWEval to evaluate whether video generators reproduce distributions of physical behaviors under identical observations and actions; current systems fail probabilistic alignment.
Reasoning
The paper provides a clear formalization and a large benchmark with 50 scenarios and 11 systems, making a convincing empirical case that current video generators are not probabilistically aligned. However, the abstract only sketches interventions to improve alignment and does not report detailed results, so the practical path forward remains unclear.
Read-first score
Read-first score 43.8, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 49.
Field roles
Rank sensitivity
Stability: volatile; rank range: 385.
Keyword Scores
Deep Analysis
Innovations
- Formalizes probabilistic alignment as a distributional criterion for world models.
- Introduces PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics.
- Introduces PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors.
Methodology
The paper formalizes probabilistic alignment for world models and introduces PAWBench with 50 scenarios. It evaluates eleven current video generation systems using the PAWEval protocol, which converts repeated video rollouts into empirical distributions over possible physical behaviors. It also tests whether language prompts, initial noise sampling, or model training can reshape predictive distributions.
Key Results
Across 50 scenarios and eleven current systems, no model consistently matches the reference probabilities while recovering the range of valid behaviors.