A Comprehensive Survey on World Models for Embodied AI
TLDR
A comprehensive survey proposing a unified framework and taxonomy for world models in embodied AI, covering functionality, temporal modeling, spatial representation, and open challenges.
Reasoning
Strengths: Provides a structured taxonomy and comprehensive overview of world models for embodied AI, including data resources, metrics, and quantitative comparisons. Weaknesses: As a survey, it does not introduce novel methods or experiments; the taxonomy may be subjective and open challenges are well-known.
Read-first score
Read-first score 55.4, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 52.
Field roles
Rank sensitivity
Stability: volatile; rank range: 833.
Keyword Scores
Deep Analysis
Innovations
- Proposes a unified framework for world models in embodied AI, formalizing problem setting and learning objectives.
- Introduces a three-axis taxonomy: Functionality (Decision-Coupled vs. General-Purpose), Temporal Modeling (Sequential Simulation and Inference vs. Global Difference Prediction), and Spatial Representation (Global Latent Vector, Token Feature Sequence, Spatial Latent Grid, Decomposed Rendering Representation).
- Systematizes data resources and metrics across robotics, autonomous driving, and general video settings, covering pixel prediction quality, state-level understanding, and task performance.
- Offers a quantitative comparison of state-of-the-art world models.
- Distills key open challenges for the field.
Methodology
This survey conducts a comprehensive literature review of world models for embodied AI, formalizing the problem and learning objectives. It constructs a three-axis taxonomy to categorize existing approaches, systematizes available datasets and evaluation metrics, and performs a quantitative comparison of state-of-the-art models to identify trends and gaps.
Key Results
The survey provides a quantitative comparison of state-of-the-art world models and distills key open challenges, including the scarcity of unified datasets, the need for physical-consistency metrics over pixel fidelity, the trade-off between performance and real-time computational efficiency, and the difficulty of achieving long-horizon temporal consistency while mitigating error accumulation.
Limitations
- Scarcity of unified datasets for training and evaluation.
- Need for evaluation metrics that assess physical consistency over pixel fidelity.
- Trade-off between model performance and computational efficiency required for real-time control.
- Core modeling difficulty of achieving long-horizon temporal consistency while mitigating error accumulation.