Awesome World Model Hub Papers · Datasets · Projects
← Back to papers

Video Generation Models in Robotics - Applications, Research Challenges, Future Directions

arXiv 2026 60.6 method

TLDR

Survey on video generation models as world models in robotics, covering applications, challenges, and future directions.

Reasoning

The paper provides a comprehensive survey of video generation models applied as world models in robotics, highlighting strengths like photorealistic simulation and fine-grained dynamics. Weaknesses include a lack of original empirical experiments and reliance on prior work, with limited discussion of practical deployment challenges.

Read-first score

Read-first score 60.6, weighted from topical fit, citation, graph, method, reproducibility, and recency signals. Original total remains 50.

Recency 8%
100

Uses a gentle age decay so recent papers surface without erasing older foundations. 2026

Topical relevance 42%
71.4

Uses existing LLM keyword relevance scores normalized to 0-100. world model,world simulator,generative world model,interactive world model,video world model,world dynamics prediction,model-based reinforcement learning world model

Methodology quality 25%
60

Screens visible abstract and analysis fields for experiment, dataset, baseline, metric, and limitation evidence. markers=evaluation

Reproducibility 25%
30

Screens links and visible text for paper, code, dataset, artifact, and repository signals. pdf=True; code=False; dataset=False; markers=none

Field roles

Frontier

Rank sensitivity

Stability: volatile; rank range: 399.

Keyword Scores

world model
9
video world model
9
world dynamics prediction
8
generative world model
7
model-based reinforcement learning world model
7
world simulator
6
interactive world model
4

Deep Analysis

Innovations

  • Comprehensive survey of video generation models as embodied world models in robotics, covering applications across imitation learning, reinforcement learning, visual planning, and policy evaluation.
  • Identification of key challenges (poor instruction following, hallucinations, safety, high costs) that hinder trustworthy integration of video models in robotics.
  • Articulation of future research directions to address these challenges and enable broader adoption in safety-critical settings.

Methodology

This survey reviews existing literature on video generation models and their applications in robotics, categorizing them by use cases such as data generation, action prediction, dynamics modeling, planning, and evaluation, and then synthesizes current challenges and future directions.

Key Results

The survey finds that video models are being used for photorealistic simulation, world modeling, and policy learning, but face major hurdles including poor instruction following, physical violations (hallucinations), unsafe content, and high computational costs.

Limitations

  • Poor instruction following
  • Hallucinations and violations of physics
  • Unsafe content generation
  • Significant data curation, training, and inference costs

Tags