World Models for Robotic Manipulation: A Survey
TLDR
A survey of world models for robotic manipulation, defining them as action-conditioned predictive systems and organizing approaches by representation, prediction-action coupling, and usage pipeline.
Reasoning
The paper provides a comprehensive taxonomy and review of world models in robotic manipulation, clearly defining the scope and distinguishing different families. However, as a survey, it lacks novel experiments or real-world validation, and the breadth may obscure specific design trade-offs.
Read-first score
Read-first score 64, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 51.
Field roles
Rank sensitivity
Stability: volatile; rank range: 396.
Keyword Scores
Deep Analysis
Innovations
- Operational definition of world model as action-conditioned predictive system
- Organization of existing work into five representation families
- Functional taxonomy separating integrated prediction-action models from explicit predictive planners
- Characterization of infrastructure roles (synthetic experience generation, candidate filtering, search-based evaluation, learned environments, outcome verification)
- Mapping of roles across pretraining, post-training, and inference adaptation
- Review of 34 manipulation datasets and synthesis of evaluation protocols
Methodology
The authors conduct a systematic survey of world models for robotic manipulation, organizing the literature by three guiding questions: what future representation is predicted, how prediction is connected to action, and when prediction is used. They develop a functional taxonomy, characterize infrastructure roles, review 34 datasets, and synthesize evaluation protocols for predictive fidelity, task performance, and simulator reliability.
Key Results
The survey reveals that world models are evolving from task-specific dynamics predictors into predictive infrastructure for robot learning, and identifies critical open challenges including contact modeling, hallucination control, action alignment, and benchmarking under closed-loop use.
Limitations
- Contact modeling remains an open challenge
- Hallucination control is not yet solved
- Action alignment between prediction and control is difficult
- Benchmarking under closed-loop use is underdeveloped