MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction
MaskGWM combines diffusion transformers with MAE-style mask reconstruction for generalizable driving world models, enabling long-horizon and multi-view video prediction.
Research papers, datasets, and open-source projects
Automatically curated from research paper sources. Built with awesome-hub-generator.
Categories: application · benchmark · method · survey · system · theory
Latest papers
MaskGWM combines diffusion transformers with MAE-style mask reconstruction for generalizable driving world models, enabling long-horizon and multi-view video prediction.
Proposes LoopNav, a Minecraft dataset and benchmark for evaluating spatial consistency in world models using loop-based navigation.
Proposes cRSSM, a contextual world model for Dreamer, improving zero-shot generalization to unseen dynamics in contextual RL.
STORM combines Transformers and VAEs for efficient world models in RL, achieving 126.7% human performance on Atari 100k with fast training.
RoboDreamer learns compositional world models by factorizing video generation from language, enabling generalization to unseen tasks and goals.
RoboScape is a physics-informed embodied world model that jointly learns RGB video generation and physics knowledge for realistic robotic video synthesis.
Datasets

Proposes LoopNav, a Minecraft dataset and benchmark for evaluating spatial consistency in world models using loop-based navigation.


An open-source benchmark and baseline model for evaluating action fidelity in world models for autonomous driving.
Projects
GitHub
Collect some World Models for Autonomous Driving (and Robotic, etc.) papers.