Ctrl-World: A Controllable Generative World Model for Robot Manipulation
TLDR
A controllable multi-view generative world model for robot manipulation, enabling policy evaluation and improvement via long-horizon consistent imagination.
Reasoning
The paper addresses a key challenge in building controllable world models for generalist robot policies, using a large real-world dataset and achieving long-horizon consistency. However, the abstract is cut off, and potential limitations include reliance on a single dataset and lack of real-world deployment validation.
Read-first score
Read-first score 71.1, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 62.
Field roles
Rank sensitivity
Stability: volatile; rank range: 174.
Keyword Scores
Deep Analysis
Innovations
- Controllable multi-view world model for evaluating and improving generalist robot policies
- Pose-conditioned memory retrieval mechanism for long-horizon consistency
- Frame-level action conditioning for precise action control
Methodology
The model is trained on the DROID dataset (95k trajectories, 564 scenes) and generates spatially and temporally consistent trajectories for over 20 seconds. It uses a pose-conditioned memory retrieval mechanism to maintain long-horizon consistency and frame-level action conditioning to achieve precise action control, enabling multi-step interactions with generalist robot policies.
Key Results
The method can accurately rank policy performance without real-world robot rollouts, and by synthesizing successful trajectories in imagination for supervised fine-tuning, it improves policy success by 44.7%.
Limitations
- Evaluation is limited to the DROID dataset, which may not cover all real-world scenarios
- Long-horizon consistency is demonstrated for over 20 seconds, but longer interactions may not be supported
- The approach requires a large dataset (95k trajectories) and may not generalize to novel camera placements or scenes beyond those in the training data