Adapting Vision-Language Models for Evaluating World Models
TLDR
Adapts VLMs to evaluate world model rollouts via action/character recognition tasks, achieving human-aligned results with a lightweight evaluator.
Reasoning
The paper addresses a critical gap in evaluating world model rollouts by introducing a VLM-based evaluator (UNIVERSE) with extensive experiments and human studies. Strengths include a novel evaluation protocol and strong empirical validation, but the focus is limited to two recognition tasks, potentially missing broader evaluation aspects.
Read-first score
Read-first score 57, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 34.
Field roles
Rank sensitivity
Stability: volatile; rank range: 308.
Keyword Scores
Deep Analysis
Innovations
- Adapting Vision-Language Models for fine-grained, temporally grounded evaluation of world model rollouts
- UNIVERSE: a unified VLM-based evaluator for video world model rollouts adapted under data and compute constraints
- Evaluation protocol targeting action recognition and character recognition across binary, multiple-choice, and open-ended formats
Methodology
The paper introduces an evaluation protocol with two recognition tasks (action and character) assessed in binary, multiple-choice, and open-ended formats. It presents UNIVERSE, a VLM-based evaluator adapted under data and compute constraints, and conducts extensive experiments (over 5,154 GPU-days) exploring full, partial, and parameter-efficient adaptation methods across various task formats, context lengths, sampling methods, and data compositions.
Key Results
The unified evaluator achieves parity with task-specific checkpoints, and human studies across seven diverse environments confirm strong alignment with human judgments.