Evaluating Gemini Robotics Policies in a Veo World Simulator
TLDR
Demonstrates using a frontier video model (Veo) as a generative world simulator for comprehensive evaluation of robotics policies, including nominal, OOD, and safety scenarios.
Reasoning
Strengths: novel comprehensive evaluation framework using video models, enabling OOD and safety testing. Weaknesses: limited to one video model, and evaluation is simulated rather than real-world; abstract lacks details on quantitative results.
Read-first score
Read-first score 65.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 56.
Field roles
Rank sensitivity
Stability: volatile; rank range: 295.
Keyword Scores
Deep Analysis
Innovations
- Using video models for the entire spectrum of policy evaluation: nominal performance, out-of-distribution generalization, and physical/semantic safety probing.
- A generative evaluation system built on a frontier video foundation model (Veo) optimized for robot action conditioning and multi-view consistency.
- Integration of generative image-editing and multi-view completion to synthesize realistic scene variations along multiple axes of generalization.
- Demonstration that the system can predict relative policy performance, determine impact of generalization axes, and perform red teaming for safety violations.
Methodology
The system is built upon a frontier video foundation model (Veo), optimized for robot action conditioning and multi-view consistency. It integrates generative image-editing and multi-view completion to synthesize realistic variations of real-world scenes. The system is evaluated with 1600+ real-world evaluations of eight Gemini Robotics policy checkpoints across five tasks for a bimanual manipulator.
Key Results
The system preserves base video model capabilities to accurately simulate edited scenes, enabling accurate prediction of relative policy performance in nominal and OOD conditions, determination of generalization axes impact, and red teaming for safety violations.
Limitations
- The system's performance is contingent on the capabilities of the underlying Veo video model.
- The evaluation is limited to eight policy checkpoints and five tasks for a bimanual manipulator.