Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Evaluating Gemini Robotics Policies in a Veo World Simulator

arXiv 25.12 2025 65.1 system, application

TLDR

Demonstrates using a frontier video model (Veo) as a generative world simulator for comprehensive evaluation of robotics policies, including nominal, OOD, and safety scenarios.

Reasoning

Strengths: novel comprehensive evaluation framework using video models, enabling OOD and safety testing. Weaknesses: limited to one video model, and evaluation is simulated rather than real-world; abstract lacks details on quantitative results.

Read-first score

Read-first score 65.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 56.

Recency 8%
86.7

Uses a gentle age decay so recent papers surface without erasing older foundations. 2025

Topical relevance 42%
80

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=evaluation

Reproducibility 25%
38

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=checkpoint

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 295.

Keyword Scores

world model
9
world simulator
9
generative world model
9
video world model
9
interactive world model
8
world dynamics prediction
7
model-based reinforcement learning world model
5

Deep Analysis

Innovations

  • Using video models for the entire spectrum of policy evaluation: nominal performance, out-of-distribution generalization, and physical/semantic safety probing.
  • A generative evaluation system built on a frontier video foundation model (Veo) optimized for robot action conditioning and multi-view consistency.
  • Integration of generative image-editing and multi-view completion to synthesize realistic scene variations along multiple axes of generalization.
  • Demonstration that the system can predict relative policy performance, determine impact of generalization axes, and perform red teaming for safety violations.

Methodology

The system is built upon a frontier video foundation model (Veo), optimized for robot action conditioning and multi-view consistency. It integrates generative image-editing and multi-view completion to synthesize realistic variations of real-world scenes. The system is evaluated with 1600+ real-world evaluations of eight Gemini Robotics policy checkpoints across five tasks for a bimanual manipulator.

Key Results

The system preserves base video model capabilities to accurately simulate edited scenes, enabling accurate prediction of relative policy performance in nominal and OOD conditions, determination of generalization axes impact, and red teaming for safety violations.

Limitations

  • The system's performance is contingent on the capabilities of the underlying Veo video model.
  • The evaluation is limited to eight policy checkpoints and five tasks for a bimanual manipulator.

Tags