MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction
MaskGWM combines diffusion transformers with MAE-style mask reconstruction for generalizable driving world models, enabling long-horizon and multi-view video prediction.
Research index
Automatically collected and categorized papers from arXiv, upstream awesome lists, and other research sources.
MaskGWM combines diffusion transformers with MAE-style mask reconstruction for generalizable driving world models, enabling long-horizon and multi-view video prediction.
Proposes LoopNav, a Minecraft dataset and benchmark for evaluating spatial consistency in world models using loop-based navigation.
Proposes cRSSM, a contextual world model for Dreamer, improving zero-shot generalization to unseen dynamics in contextual RL.
STORM combines Transformers and VAEs for efficient world models in RL, achieving 126.7% human performance on Atari 100k with fast training.
RoboDreamer learns compositional world models by factorizing video generation from language, enabling generalization to unseen tasks and goals.
RoboScape is a physics-informed embodied world model that jointly learns RGB video generation and physics knowledge for realistic robotic video synthesis.
UWM couples video and action diffusion in a unified transformer for pretraining on large robotic datasets, enabling policy and dynamics learning.
MIND is a benchmark for evaluating memory consistency and action control in world models using 250 high-quality videos across diverse scenes and action spaces.
DIAMOND uses diffusion models for world modeling in Atari, achieving state-of-the-art agent performance with improved visual details.
Proposes Parallel Observation Prediction to accelerate token-based world models, achieving 15.4x faster imagination and superhuman Atari performance.
IRASim is a fine-grained world model for robot manipulation that generates videos conditioned on actions, improving action-frame alignment and enabling planning.
An open-source benchmark and baseline model for evaluating action fidelity in world models for autonomous driving.
Vista is a generalizable driving world model with high fidelity and versatile controllability, outperforming prior methods on multiple datasets.
MineWorld is a real-time interactive world model for Minecraft using a visual-action autoregressive Transformer with parallel decoding, outperforming diffusion-based models.
Astra is an interactive general world model using autoregressive denoising for long-term video prediction with action control across diverse real-world tasks.
Drive-WM: a driving world model that generates multiview videos for safe planning by forecasting multiple futures and selecting optimal trajectories via image-based rewards.
Proposes Event-Aware World Model (EAWM) that learns event representations from raw observations to improve MBRL generalization and robustness.
Yume generates interactive, dynamic worlds from images using keyboard control, with a framework including camera motion quantization, MVDT, and advanced sampling.
AVID adapts pretrained video diffusion models to action-conditioned world models using a learned mask, without accessing model parameters.
A survey on interactive video world modeling, covering frontiers, challenges, benchmarks, and future trends.
FAR-Drive proposes a frame-level autoregressive video generation framework for closed-loop autonomous driving simulation, achieving state-of-the-art on nuScenes.
Proposes Dynamic World Simulation (DWS) to turn pre-trained video generative models into controllable world simulators using action-conditioned modules and motion-reinforced loss.
A scalable generative model for autonomous driving that learns scene dynamics, enabling simulation, scenario generation, and planning, achieving SOTA on real-world benchmarks.
BEVWorld transforms multimodal sensor inputs into unified BEV latent space for autonomous driving, using a tokenizer and diffusion model to forecast future scenes.
Epona is an autoregressive diffusion world model for autonomous driving enabling long-horizon video prediction and trajectory planning with state-of-the-art performance.
GeoDrive integrates 3D geometry into driving world models for precise action control, improving spatial awareness and scene modeling.
LongScape combines intra-chunk diffusion and inter-chunk autoregression with action-guided chunking and Context-aware MoE for stable long-horizon video generation in embodied world models.
A transformer-based world model (TWM) achieves sample-efficient RL on Atari 100k by autoregressively processing states, actions, and rewards.
DrivingGen is the first comprehensive benchmark for generative driving world models, with new metrics and diverse data to evaluate visual realism, trajectory plausibility, temporal coherence, and controllability.
GenRL uses multimodal foundation world models to connect VLMs with generative world models for RL, enabling task specification via vision/language and multi-task generalization.
Examines test-time verifiers for world-model-based spatial reasoning, finding pitfalls and introducing ViSA, which improves on SAT-Real but not MMSI-Bench.
R2-Dreamer introduces a redundancy-reduction objective for decoder-free MBRL, achieving faster training and competitive performance without data augmentation.
ADriver-I is a general world model for autonomous driving using interleaved vision-action pairs, MLLM, and diffusion to predict control signals and future frames.
A comprehensive survey of general world models, focusing on Sora's simulation capabilities, generative video, autonomous driving, and autonomous agents.
WorldDrive unifies vision and motion representation in a driving world model to couple scene generation and real-time planning.
UniFuture is a unified 4D driving world model that jointly generates future RGB and depth sequences, outperforming specialized models on nuScenes and Waymo.
KeyWorld uses key frame reasoning to accelerate robotic world models by focusing computation on semantic key frames, achieving 5.68x speedup on LIBERO.
Dynalang learns a multimodal world model from diverse language to predict future observations and rewards, enabling game-playing and navigation in photorealistic homes.
Introduces LingBot-VA, an autoregressive diffusion framework for video world modeling and robot control, evaluated in simulation and real-world.
UniWM integrates visual foresight and planning in a unified memory-augmented world model for visual navigation, achieving up to 30% improvement on benchmarks.
HMA uses heterogeneous masked autoregression to model action-video dynamics for robot learning, achieving faster and more controllable video generation.
LiveWorld extends video world models to simulate persistent out-of-sight dynamics using a global state and monitor-based mechanism, evaluated on LiveBench.
OccSora is a diffusion-based 4D occupancy generation model that simulates driving scenes as a world simulator for autonomous driving.
LatentDriver uses a latent world model with mixture distributions to handle uncertainty and self-delusion in autonomous driving, outperforming SOTA on Waymax.
Matrix-Game 2.0 is an open-source interactive world model for real-time streaming video generation using few-step auto-regressive diffusion at 25 FPS.
LingBot-World is an open-source world simulator from video generation with high fidelity, long-term memory, and real-time interactivity.
minWM is an open-source framework for building real-time interactive video world models from video diffusion models via controllable fine-tuning and distillation.
DINO-world uses DINOv2 latent space to train a video world model that predicts future frames and outperforms prior models on benchmarks.