World4Omni: A Zero-Shot Framework from Image Generation World Model to Robotic Manipulation
TLDR
A zero-shot framework using image-generative VLMs as world models to generate goal states for generalizable robotic manipulation.
Reasoning
The paper presents a novel approach leveraging image-generative VLMs as world models for zero-shot robotic manipulation, with real-world experiments demonstrating strong performance. However, the abstract lacks details on limitations and the scope is narrow, focusing only on goal state generation.
Read-first score
Read-first score 40.5, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 19.
Field roles
Rank sensitivity
Stability: volatile; rank range: 280.
Keyword Scores
Deep Analysis
Innovations
- Leveraging Image-Generative VLMs as world models to generate desired goal states for robotic manipulation
- Using object state representation as a golden interface to separate high-level and low-level policies, enabling training-free low-level control
- Introducing a Reflection-through-Synthesis process that iteratively validates and refines the generated goal image before execution
Methodology
Goal-VLA is a zero-shot framework that uses Image-Generative VLMs as world models to generate desired goal states. From these generated images, the target object pose is derived, which serves as spatial cues for training-free low-level control. The system is evaluated in both simulated and real-world environments.
Key Results
The framework achieves strong performance and inspiring generalizability in manipulation tasks across simulated and real-world experiments.