Awesome World Model Hub Papers · Datasets · Projects

Dataset zoo

Datasets

Datasets grouped by representation and task.

Preview for IRASim
2024

IRASim

IRASim is a fine-grained world model for robot manipulation that generates videos conditioned on actions, improving action-frame alignment and enabling planning.

Preview for Towards Interactive Video World Modeling
2026

Towards Interactive Video World Modeling

A comprehensive survey on interactive world modeling, covering trends, challenges, benchmarks, and future directions in action-conditioned video/3D generation.

world modelinginteractive video generationaction-conditioned generationsurveybenchmarksCV
Preview for DrivingGen
2026

DrivingGen

DrivingGen is the first comprehensive benchmark for generative driving world models, with new metrics and diverse data to evaluate visual realism, trajectory plausibility, temporal coherence, and controllability.

Preview for ReWorld
2026

ReWorld

ReWorld uses reinforcement learning to align video-based embodied world models with physical realism, task logic, and visual quality via multi-dimensional reward modeling.

Preview for Solaris
2026

Solaris

Solaris introduces a multiplayer video world model for Minecraft, enabling consistent multi-view simulation of multi-agent interactions.

Preview for OSCAR
2026

OSCAR

OSCAR is an action-conditioned video world model for robotics that generalizes across embodiments and enables virtual policy evaluation with real-world correlation.

roboticsworld modelvideo predictionembodimentaction conditioninggeneralization
Preview for WBench
2026

WBench

WBench is a multi-turn benchmark evaluating interactive video world models across five dimensions with 289 test cases and 1,058 interaction turns.

interactive world modelsbenchmarkmulti-turn evaluationvideo qualityphysics complianceinteraction adherence
Preview for Omni-WorldBench
2026

Omni-WorldBench

Omni-WorldBench evaluates world models' interactive response in 4D settings using a prompt suite and agent-based metrics, revealing limitations across 18 models.

Preview for WHALE
2024

WHALE

WHALE introduces behavior-conditioning and retracing-rollout to improve generalizability and uncertainty estimation in world models for embodied decision-making.

Preview for Physics-IQ Verified
2026

Physics-IQ Verified

Audits and improves Physics-IQ benchmark for evaluating physical understanding of video generative models, refining samples and prompts.

video generationworld modelingphysical understandingbenchmark evaluationvideo generative modelsCV
Preview for ARB4WM
2026

ARB4WM

ARB4WM is a unified benchmark for evaluating adversarial robustness of world models in continuous control across policy, value, and latent-dynamics levels.

adversarial robustnessworld modelscontinuous controlbenchmarkvisual perturbationsAI
Preview for MIND
2026

MIND

MIND is a benchmark for evaluating memory consistency and action control in world models using 250 high-quality videos across diverse scenes and action spaces.

Preview for Qwen-RobotWorld Technical Report
2026

Qwen-RobotWorld Technical Report

A language-conditioned video world model unifying embodied intelligence across robotics, driving, navigation, and human-to-robot transfer.

embodied AIworld modelvideo generationlanguage-conditionedroboticsautonomous driving
Preview for stable-worldmodel
2026

stable-worldmodel

An open-source platform for standardized, reproducible world modeling research with high-performance data handling, baseline implementations, and systematic evaluation benchmarks.

world modelsreproducibilityevaluationopen-sourcedata pipelinegeneralization
Preview for EWMBench
2025

EWMBench

Proposes EWMBench, a benchmark evaluating embodied world models on scene consistency, motion correctness, and semantic alignment using a curated dataset and multi-dimensional toolkit.

Preview for WorldModelBench
2025

WorldModelBench

Proposes WorldModelBench to evaluate video generation models as world models, focusing on physics adherence and instruction-following with human labels.

Preview for SimWorld
2025

SimWorld

Proposes SimWorld, a benchmark combining simulation engine and world model for controllable scene generation to improve autonomous driving perception models.

Preview for WorldOlympiad
2026

WorldOlympiad

A benchmark evaluating video-based world models across physical, geometric, and interaction fidelity with real-world scenarios.

world modelsvideo generationbenchmarkphysical faithfulnessgeometric consistencyinteraction fidelity
Preview for Dexterous World Models
2025

Dexterous World Models

A video diffusion framework that models dexterous human actions inducing dynamic changes in static 3D scenes for interactive digital twins.

Preview for WorldReasonBench
2026

WorldReasonBench

Introduces WorldReasonBench, a benchmark to evaluate video generators as world-state predictors, revealing a gap between visual quality and reasoning.

video generationbenchmarkworld-state predictionreasoningconsistencyevaluation
Preview for A Survey
2025

A Survey

A survey on integrating physical simulators and world models for embodied intelligence, bridging simulation and real-world deployment.

Preview for EchoWorld
2025

EchoWorld

EchoWorld uses motion-aware world models for echocardiography probe guidance, reducing errors via pre-trained masked prediction and motion-aware attention.

Preview for Holo-World
2026

Holo-World

Holo-World enables unified camera, object, and weather control from a single image for video generation.

video world modelcontrollable video generationcamera controlobject controlweather transferdataset
Preview for What-If World
2026

What-If World

Introduces What-If World, a causal benchmark using prompt pairs to test if video world models correctly respond to physical changes.

video generationworld modelscausal reasoningbenchmarkphysical plausibilityembodied scenarios
Preview for Current World Models Lack a Persistent State Core
2026

Current World Models Lack a Persistent State Core

This paper introduces WRBench to test if world models maintain persistent internal state when unobserved, finding current models fail to evolve events during occlusion.

world modelsbenchmarkpersistent statecamera motionobservabilityCV
Preview for VerseCrafter
2026

VerseCrafter

VerseCrafter uses 4D geometric control (point clouds and 3D Gaussian trajectories) to generate realistic, view-consistent videos with precise camera and multi-object motion.

Preview for PointWorld
2026

PointWorld

PointWorld is a large pre-trained 3D world model that forecasts 3D point flows from RGB-D images and actions, enabling real-time MPC for robotic manipulation across embodiments.

Preview for World Models for Robotic Manipulation
2026

World Models for Robotic Manipulation

Survey of world models for robotic manipulation, categorizing representations, action connections, and usage in robot learning pipelines.

world modelsrobotic manipulationlatent dynamicsaction-conditioned video generationphysics-informed simulationRO
Preview for Toward Stable World Models
2025

Toward Stable World Models

Introduces world stability metric for diffusion-based world models, evaluates state-of-the-art models, and proposes improvement strategies.

Preview for WorldFly
2026

WorldFly

WorldFly integrates world models into a VLA framework for UAV navigation, using dual-branch flow matching to predict future video and actions.

UAV navigationVision-Language-ActionWorld ModelUrban CanyonNavigationAI
Preview for Is Your Driving World Model an All-Around Player?
2026

Is Your Driving World Model an All-Around Player?

WorldLens benchmark evaluates driving world models across pixel quality, geometry, closed-loop driving, and human perception, revealing no model excels universally.

driving world modelsbenchmarkevaluationautonomous drivingvideo generationrealism
Preview for WorldLens
2025

WorldLens

WorldLens is a full-spectrum benchmark for evaluating driving world models across visual, geometric, physical, and behavioral fidelity.

Preview for ActWorld
2026

ActWorld

ActWorld extends navigation-centric world models to support object interaction via action-aware memory and a new dataset.

interactive world modelsaction-aware memoryobject interactionnavigationchunk-autoregressiveCV
Preview for Embody4D
2026

Embody4D

Embody4D transforms monocular robot videos into novel-view videos for embodied 4D world modeling.

embodied AIworld modelnovel view synthesisvideo generation4D modelingdata engine
Preview for KAN-Dreamer
2025

KAN-Dreamer

Integrates KANs into DreamerV3 world model, achieving parity in sample efficiency and training speed on a control task.

Preview for GigaWorld-Policy
2026

GigaWorld-Policy

GigaWorld-Policy is an action-centered World-Action Model that efficiently predicts future actions and optionally generates videos for robot policy learning.

Preview for CityBench
2024

CityBench

CityBench is a systematic benchmark using an interactive simulator to evaluate LLMs and VLMs as world models for diverse urban tasks across 13 cities.

Preview for OccDirector
2026

OccDirector

OccDirector generates 4D occupancy dynamics from natural language for autonomous driving simulation, achieving state-of-the-art instruction-following.

4D occupancyautonomous drivinglanguage-guided generationmulti-agent interactionworld modelvideo generation
Preview for VideoWorldmodel/Evaluation
2026

VideoWorldmodel/Evaluation

Matrix-Game 2.0 Self-Consistency Reproduction This README explains how to reproduce the Matrix-Game 2.0 self-consistency (SC) row in Reasoning-Structured Videos: A Stratified Diagnostic Suite for Compositional Consistency in World Models. All experimental data, generated videos, and evaluation files required for this reproduction are available on the current dataset page under Files and versions: https://huggingface.co/datasets/VideoWorldmodel/Evaluation No external dataset or… See the full description on the dataset page: https://huggingface.co/datasets/VideoWorldmodel/Evaluation.

region:usworld-modelvideo-generationbenchmarkreproducibility
Preview for WorldSimBench
2024

WorldSimBench

WorldSimBench proposes a dual evaluation framework for video generation models as world simulators, covering embodied scenarios.

Preview for Prisma-World
2026

Prisma-World

Prisma-World generates consistent multi-agent videos from multiple camera viewpoints using geometry-aware denoising and cross-view attention.

video world modelsmulti-agentcamera controlcross-view consistencygeometry-aware denoisingCV
Preview for DrivingDojo Dataset
2024

DrivingDojo Dataset

DrivingDojo dataset enables interactive world models with diverse driving maneuvers and action-controlled future prediction benchmark.

Preview for FieldSeer I
2025

FieldSeer I

Geometry-aware world model forecasts electromagnetic field dynamics from partial observations, enabling interactive digital twins for photonic design.

Preview for 3D-VLA
2024

3D-VLA

3D-VLA integrates 3D perception, reasoning, and action via a generative world model using LLM and diffusion models, improving embodied planning.

Preview for WorldMark
2026

WorldMark

WorldMark is a unified benchmark for interactive video world models, enabling fair comparison via standardized scenes, actions, and evaluation metrics.

benchmarkworld modelsinteractive video generationevaluationaction mappingcomputer vision
Preview for From Word to World
2025

From Word to World

LLMs can serve as implicit text-based world models for agentic RL, but benefits depend on behavioral coverage and environment complexity.

Preview for OmniWorld
2025

OmniWorld

OmniWorld is a large-scale, multi-domain, multi-modal dataset for 4D world modeling, enabling benchmarks and improving SOTA in reconstruction and video generation.

Preview for PhysicsMind
2026

PhysicsMind

PhysicsMind is a unified benchmark with real and simulated environments to evaluate physical reasoning and prediction in VLMs and world models.

Preview for NU-World-Model-Embodied-AI/phyground
2026

NU-World-Model-Embodied-AI/phyground

PhyGround Project Page | GitHub | Paper PhyGround is a criteria-grounded benchmark for evaluating physical reasoning in video generation. The benchmark contains 250 curated prompts, each augmented with an expected physical outcome, and a taxonomy of 13 physical laws across solid-body mechanics, fluid dynamics, and optics. Sample Usage You can download the benchmark prompts and first-frame images using the Hugging Face CLI: huggingface-cli download --repo-type dataset… See the full description on the dataset page: https://huggingface.co/datasets/NU-World-Model-Embodied-AI/phyground.

task_categories:text-to-videotask_categories:image-to-videosize_categories:n<1Kformat:jsonmodality:textmodality:video
Preview for DeTrack
2026

DeTrack

A benchmark and altitude-aware dual world model for drone-embodied tracking in interactive 3D environments.

drone-embodied trackingaerial object trackingbenchmarkactive perceptionclosed-loop controldual world model
Preview for SmallWorlds
2025

SmallWorlds

Introduces SmallWorld Benchmark for systematically evaluating world models' dynamics understanding in isolated, controlled environments.

Preview for Unified World Models
2025

Unified World Models

UWM couples video and action diffusion in a unified transformer for pretraining on large robotic datasets, enabling policy and dynamics learning.

Preview for SlowFast-VGen
2024

SlowFast-VGen

SlowFast-VGen introduces dual-speed learning combining slow world dynamics and fast episodic memory for action-driven long video generation.

Preview for ReactSim-Bench
2026

ReactSim-Bench

Introduces ReactSim-Bench, a benchmark for evaluating reactive behavior world model simulation in autonomous driving with decoupled control and safety metrics.

autonomous drivingbehavior simulationreactive capabilitybenchmarkingworld modelRO
Preview for 4DWorldBench
2025

4DWorldBench

4DWorldBench is a unified evaluation framework for 3D/4D world generation models, assessing perceptual quality, alignment, physical realism, and consistency.

Preview for Dreamland
2025

Dreamland

Dreamland combines physics simulator and generative models for controllable, photorealistic world creation, improving image quality and controllability.

Preview for PRISM
2026

PRISM

PRISM extracts action priors from a world model's encoder to guide sampling in model-based planning, improving success rates in continuous control tasks.

world modelsmodel-based planningaction priorcontinuous controlreinforcement learningimagination sampling
Preview for Cam4DOCC
2024

Cam4DOCC

Introduces Cam4DOcc, a benchmark for camera-only 4D occupancy forecasting in autonomous driving, using multiple datasets and baselines.

Preview for NarrativeWorldBench
2026

NarrativeWorldBench

Introduces a benchmark and latent world model for long-horizon co-creative audio drama, outperforming frontier LLMs in consistency and controllability.

audio dramalong-horizonbenchmarknarrative understandingstate-space modelLLM evaluation
Preview for EgoCS-400K
2026

EgoCS-400K

EgoCS-400K is a large-scale egocentric Counter-Strike dataset with temporally aligned video-action-language trajectories for training interactive world models.

egocentricdatasetworld modelsCounter-StrikegameplayCV
Preview for WorldArena
2026

WorldArena

WorldArena benchmark evaluates embodied world models on perception and functional utility, revealing a gap between visual quality and task capability.

Preview for YoCausal
2026

YoCausal

YoCausal benchmarks video diffusion models on causality using real-world counterfactual videos, revealing a gap between temporal perception and causal understanding.

video diffusion modelsworld modelscausalitybenchmarkcounterfactual reasoningvideo generation
Preview for Text2World
2025

Text2World

Introduces Text2World, a PDDL-based benchmark for evaluating LLMs on symbolic world model generation, revealing limited capabilities despite RL-trained reasoning models.

Preview for LEIA
2026

LEIA

LEIA is a world model for interactive simulation of architected materials, enabling real-time deformation and stress field prediction.

world modelarchitected materialsmachine learninginteractive simulationfinite element methodmicrostructure
Preview for ReflectiChain
2026

ReflectiChain

ReflectiChain bridges LLM and RL gaps with a generative supply chain world model and double-loop learning, improving resilience under uncertainty.

supply chain resilienceepistemic groundingworld modellarge language modelsreinforcement learninguncertainty quantification
Preview for World-in-World
2025

World-in-World

Introduces World-in-World, a platform for benchmarking world models in closed-loop embodied environments, revealing that controllability and post-training scaling matter more than visual quality.

Preview for Teaching Video Generators to Remember
2026

Teaching Video Generators to Remember

ReMind elicits dynamic memory in video generators via memory-oriented data and curriculum training, achieving state-of-the-art on STEVO-Bench.

video generationworld modelsdynamic memorydiffusion transformersstate evolutionout-of-sight reasoning
Preview for MBench
2026

MBench

MBench is a benchmark evaluating memory capability of video world models via entity, environment, and causal consistency using real-captured long videos.

video world modelsmemory capabilitybenchmarkevaluationconsistencyCV
Preview for On Memory
2026

On Memory

Compares memory mechanisms in transformer-based world models to extend memory span and reduce perceptual drift in long rollouts.

Preview for WorldArena 2.0
2026

WorldArena 2.0

WorldArena 2.0 expands embodied world model benchmarking across modality, functionality, and platform, including real-world robotic evaluations.

embodied world modelsbenchmarkmultimodalvisuotactileroboticscomputer vision
Preview for JEDI
2025

JEDI

Proposes JEDI, a latent diffusion world model to address performance asymmetry in MBRL on Atari100k, achieving balanced human-normalized scores.

Preview for NORA-1.5
2025

NORA-1.5

NORA-1.5 enhances VLA models with flow-matching action expert and world model-based preference rewards, improving reliability in simulation and real-world tasks.

Preview for TD-MPC2
2024

TD-MPC2

TD-MPC2 improves model-based RL with scalable, robust world models, achieving strong results across 104 tasks with a single hyperparameter set.

Preview for GigaWorld-1
2026

GigaWorld-1

Systematic study of world models for robot policy evaluation using WMBench benchmark, analyzing 7 video world models and 324k rollouts.

Preview for How Mobile World Model Guides GUI Agents?
2026

How Mobile World Model Guides GUI Agents?

Investigates how mobile world models guide GUI agents by comparing four modalities and evaluating downstream utility on benchmarks.

GUI agentsworld modelsvision-language modelsmobile agentsaction predictiontest-time guidance
Preview for WorldBench
2026

WorldBench

WorldBench is a video benchmark for disentangled evaluation of world models' understanding of individual physics concepts, revealing failures in SOTA models.

Preview for SimGen
2024

SimGen

Controllable synthetic data generation can substantially lower the annotation cost of training data.

Preview for Benchmarking World-Model Learning
2025

Benchmarking World-Model Learning

Proposes WorldTest protocol and AutumnBench benchmark to evaluate world models on multiple environment-level queries, showing humans outperform frontier models.

Preview for SIMMER
2026

SIMMER

SIMMER benchmarks latent failures in LLM planning using a symbolic world model, revealing high failure rates and showing counterfactual reasoning reduces them.

LLM planninglatent failuresbenchmarkworld modelautonomous agentsCL
Preview for World-R1
2026

World-R1

World-R1 uses reinforcement learning to align text-to-video generation with 3D constraints, improving geometric consistency without architectural changes.

text-to-video generation3D constraintsreinforcement learningFlow-GRPOgeometric consistencyworld simulation
Preview for IPR-1
2025

IPR-1

IPR-1 combines world-model rollouts with a VLM policy and PhysCode action space to improve physical reasoning across 1000+ games, outperforming GPT-5.

Preview for UniOcc
2025

UniOcc

UniOcc is a unified benchmark for occupancy forecasting and prediction in autonomous driving, integrating real-world and simulated data with novel metrics.

Preview for ST-Gen4D
2026

ST-Gen4D

Proposes ST-Gen4D, a world model framework for 4D generation that integrates spatiotemporal cognition via global and local graphs.

4D generationspatiotemporal cognitionworld modelgenerative modelsdynamic topologyCV
Preview for Drive&Gen
2025

Drive&Gen

Proposes DriveGen to co-evaluate end-to-end driving and video generation models using statistical measures and synthetic data.

Preview for MagicTime
2024

MagicTime

MagicTime generates time-lapse videos encoding real-world physics via decoupled training and a specialized dataset, acting as metamorphic simulators.

Preview for Vega
2026

Vega

Vision-language-action models have reshaped autonomous driving to incorporate languages into the decision-making process.

Preview for ShareVerse
2026

ShareVerse

ShareVerse enables multi-agent consistent video generation for shared world modeling using spatial concatenation and cross-agent attention on CARLA data.

Preview for Orca
2026

Orca

Orca is a general world foundation model learning unified world latent space via Next-State-Prediction from multimodal data, enabling text, image, and action generation.

Preview for PragWorld
2025

PragWorld

Evaluates LLMs' local world model robustness in conversations under minimal linguistic alterations, proposes interpretability and fine-tuning methods.

Preview for Thinking in Video
2026

Thinking in Video

Proposes Causal-Generative Dual-Judge (CGDJ) to evaluate if video generators truly reason about real-world dynamics, revealing a perception-prediction gap.

Preview for Dynamic Sparsity
2025

Dynamic Sparsity

This paper challenges common sparsity assumptions in world models for robotic RL, finding global sparsity rare but local state-dependent sparsity common.

Preview for Deform360
2026

Deform360

A large-scale multi-view visuotactile dataset for evaluating deformable object world models, comparing 2D video and 3D particle approaches.

Preview for PanoWorld
2026

PanoWorld

PanoWorld generates geometry-consistent 360° video from a single image and caption using depth and trajectory consistency losses.

panoramic video generationworld modelgeometry consistencydepth consistency360-degree videoCV
Preview for Audio-Visual World Models
2025

Audio-Visual World Models

Proposes Audio-Visual World Models (AVWM) integrating binaural audio and visual dynamics, with a benchmark and a diffusion transformer model for multimodal prediction.

Preview for RynnWorld-4D
2026

RynnWorld-4D

RynnWorld-4D generates future RGB, depth, and optical flow from a single RGB-D image and language instruction for robotic manipulation.

Preview for Apple-π
2026

Apple-π

Introduces Apple-PI, a benchmark to evaluate video generation models as law-grounded world simulators using classical mechanics tasks.

Preview for UniUGP
2025

UniUGP

UniUGP unifies understanding, generation, and planning for autonomous driving via hybrid experts and four-stage training, achieving SOTA on long-tail scenarios.

Preview for WorldDiT
2026

WorldDiT

WorldDiT is a diffusion transformer for unified action and visual world modeling, achieving strong simulation results without large VLMs.

Preview for DSWorld
2026

DSWorld

Introduces DSWorld, a data science world model for predicting execution outcomes, accelerating RL agent training and search-based inference.

Preview for EvolvingWorld
2026

EvolvingWorld

EvolvingWorld introduces an open-schema framework for co-evolving characters and world models in interactive literary simulations, with a dataset and evaluation protocol.

Preview for VideoCoCo
2026

VideoCoCo

VideoCoCo uses executable Blender code as a chain-of-thought to generate physically consistent videos via a dual-engine framework.

Preview for MemLearner
2026

MemLearner

MemLearner learns to query context memory adaptively for video world models, improving scene consistency under occlusions and dynamics.

Preview for PhysMani
2026

PhysMani

PhysMani couples a physics-principled 3D Gaussian world model with a policy for dynamic object manipulation, achieving superior success in simulation and real-world tasks.

Preview for GameFactorly
2025

GameFactorly

GameFactory generates open-domain action-controllable game videos using a multi-phase training strategy with domain adapter.

Preview for Sekai
2025

Sekai

A large-scale first-person video dataset (5000+ hours, 100+ countries) with rich annotations for training video generation models for world exploration.

Preview for CausalARC
2025

CausalARC

Introduces CausalARC, a testbed for AI reasoning using causal world models, evaluated on language models.

Preview for RetailSMV
2026

RetailSMV

Adapts a foundation video world model to retail scenes, comparing egocentric vs exocentric adaptation using a new synchronized multi-view dataset.

Preview for MemoBench
2026

MemoBench

MemoBench benchmarks world modeling in dynamic environments using a disappear-and-reappear paradigm with synthetic and real-world clips.

Preview for WorldOdysseyBench
2026

WorldOdysseyBench

Introduces WorldOdysseyBench, a benchmark for long-horizon stability of interactive world models across action, vision, physics, and memory dimensions.

Preview for Think Before You Drive
2025

Think Before You Drive

A world model-inspired framework for autonomous vehicle visual grounding that reasons about future spatial states to disambiguate natural-language commands.

Preview for OpenSTL
2023

OpenSTL

A benchmark for spatio-temporal predictive learning, comparing recurrent and recurrent-free models across multiple domains.

Preview for Large Emotional World Model
2025

Large Emotional World Model

Proposes Large Emotional World Model integrating emotion into world models, improving prediction of emotion-driven social behaviors.

Preview for VideoPhy
2024

VideoPhy

A benchmark to evaluate if text-to-video models follow physical commonsense for real-world activities, finding current models severely lacking.

Preview for VideoVerse
2025

VideoVerse

VideoVerse benchmark evaluates T2V models on temporal causality and world knowledge, revealing gaps in world model capabilities.

Preview for PAI-Bench
2025

PAI-Bench

Introduces PAI-Bench, a benchmark evaluating perception and prediction in Physical AI using 2,808 real-world cases across video tasks.

Preview for EarthNet2021
2021

EarthNet2021

EarthNet2021 dataset and challenge for forecasting satellite images conditioned on future weather, enabling high-resolution Earth surface predictions.

Preview for MetaOthello
2026

MetaOthello

Transformers trained on multiple Othello variants share a common board-state representation rather than isolating world models.

Preview for TC-Bench
2025

TC-Bench

Derived from paper: TC-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation

Preview for STDiff
2023

STDiff

Proposes STDiff, a spatio-temporal diffusion model with neural SDE for continuous stochastic video prediction, achieving state-of-the-art performance.

Preview for Thinking Ahead
2025

Thinking Ahead

Introduces Foresight Intelligence and FSU-QA dataset to evaluate VLMs and world models on reasoning about future events.

Preview for Playable Video Generation
2021

Playable Video Generation

Unsupervised learning of playable video generation where user controls video by selecting discrete actions, using self-supervised encoder-decoder with action bottleneck.

Preview for PWM-ArtGen
2026

PWM-ArtGen

A part world model for articulated 3D object generation from a single image, learning joint visual dynamics and kinematic parameters.

Preview for V-ReasonBench
2025

V-ReasonBench

Introduces V-ReasonBench, a benchmark for evaluating video reasoning in generative models across four dimensions using synthetic and real-world sequences.

Preview for Open-AoE
2026

Open-AoE

Open-AoE provides a large-scale egocentric manipulation dataset and toolchain for embodied learning, including 2000 hours of video and annotations.

Preview for WorldRoamBench
2026

WorldRoamBench

Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignore memory and interaction physics. We introduce WorldRoamBench, an open-world benchmark for long-horiz...

Preview for UniVR
2026

UniVR

UniVR uses RL to learn visual reasoning, physical dynamics, and planning from pure visual demonstrations, achieving 25% improvement on a new benchmark.

Preview for DreamTraj
2026

DreamTraj

Accurate prediction of object trajectories during manipulation is essential for closing the perception-action loop. Progress is limited on two fronts: available datasets lack fine-grained language-to-motion annotations, and existing predict...

Preview for OSCBench
2026

OSCBench

Introduces OSCBench, a benchmark evaluating object state changes in text-to-video generation; models struggle with accurate state transformations.

Preview for WildCity
2026

WildCity

WildCity is a real-world city-scale multimodal dataset and simulator for rendering, simulation, and spatial intelligence research.

Preview for ACE-Data-0
2026

ACE-Data-0

ACE-Data-0 presents a data engine capturing multimodal human-centric interactions in real homes for embodied intelligence.

Preview for RoboTrustBench
2026

RoboTrustBench

Derived from paper: RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

Preview for ChangChrisLiu/GNN_Disassembly_WorldModel
2026

ChangChrisLiu/GNN_Disassembly_WorldModel

GNN Constraint-Aware World Model Dataset (v3) Real robot episodes with per-frame constraint graphs, SAM2 segmentation masks + 256-D feature embeddings, full 3D depth bundles, and synchronized robot states across two manipulation domains. Both domains share the v3 on-disk layout (same JSON/NPZ schemas, same delta-encoded frame_states, same fully-connected PyG expansion at load time) and now share a unified 270-D node feature format — the PyG loader reads a fixed 10-D type… See the full description on the dataset page: https://huggingface.co/datasets/ChangChrisLiu/GNN_Disassembly_WorldModel.

task_categories:roboticstask_categories:image-segmentationtask_categories:graph-mllanguage:enlicense:cc-by-4.0size_categories:1K<n<10K
Preview for GroupToM-Bench
2026

GroupToM-Bench

Introduces GroupToM-Bench, a multimodal benchmark evaluating group-level theory of mind in MLLMs, revealing gaps in social world modeling.

group theory of mindmultimodal large language modelssocial emergencebenchmarkcollective behaviorCV
Preview for Kalso42/WorldModelForMaze
2026

Kalso42/WorldModelForMaze

WorldModelForMaze Code, datasets, and trained checkpoints for studying world-model representations in maze navigation, based on a modified NanoGPT. Contents *.py — training, testing, probing, and visualization scripts (see readme.md). model/ — architectures: transformer, transformer-rope, transformer-nextlat, mamba, mamba2, gated-deltanet, gru. data/maze/100/ — tokenized maze datasets for Tasks A/C/E/H/I (RWs paths, 100 nodes). out/ — final (10000-iter)… See the full description on the dataset page: https://huggingface.co/datasets/Kalso42/WorldModelForMaze.

task_categories:otherlicense:mitregion:usmazeworld-modelsequence-modeling
Preview for Kalso42/WorldModelForMazeWithX
2026

Kalso42/WorldModelForMazeWithX

WorldModelForMazeWithX Maze pathfinding sequences for training/probing sequence models (Transformer, Mamba, GRU, Gated-DeltaNet, ...). Task C1: relative-turn navigation on a fixed 10×10 directed grid. Includes a special x terminator marking wall-hit (illegal) paths, used to study a model's ability to recognize its own errors. Maze 10×10 grid, 100 nodes (0–99). Directed edges (down/right, both directions added), edge probability 0.6. Graph:… See the full description on the dataset page: https://huggingface.co/datasets/Kalso42/WorldModelForMazeWithX.

task_categories:text-generationlanguage:enlicense:mitsize_categories:1M<n<10Mregion:usmaze
Preview for kfallah/world-model-pi-sft-with-cot
2026

kfallah/world-model-pi-sft-with-cot

world-model-pi-sft-with-cot SFT-ready data for training a world model of the TeichAI pi coding-agent harness. Each world-model assistant target carries a teacher-distilled <think>{rationale}</think> block explaining why the next environment turn follows from prior context, followed by the original [developer] / [tool:*] environment block. Stage 2 of a 2-stage pipeline. Stage 1 dataset (with <COT_PLACEHOLDER> slots) is at kfallah/world-model-pi-sft-formatted. Schema… See the full description on the dataset page: https://huggingface.co/datasets/kfallah/world-model-pi-sft-with-cot.

task_categories:text-generationlanguage:enlicense:apache-2.0size_categories:1K<n<10Kregion:usworld-model
Preview for nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios
2026

nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios

Dataset Description: PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios is a large-scale synthetic video dataset of autonomous-driving scenes generated with NVIDIA's internal Omniverse simulation platform. Each clip is a temporally consistent multi-camera surround capture of one ego vehicle and surrounding traffic participants, paired with per-camera VLM captions. The dataset is designed to fill gaps in real-world driving data along two axes: (1) targeted long-tail… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios.

language:enlicense:othersize_categories:100K<n<1Mmodality:videoregion:usvideo
Preview for nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes
2026

nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes

Dataset Description: The SDG-SynHuman is a large-scale synthetic video dataset of digital humans rendered in diverse indoor and outdoor 3D environments. The dataset contains 236,937 clips, totaling approximately 5,841 hours of video, and is designed to support training and post-training of NVIDIA Cosmos world foundation models and related physical AI research. Each sample is a temporally coherent 60-120 second video clip rendered at 1080p and 30 fps. Clips contain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes.

language:enlicense:othermodality:videoregion:usvideosynthetic
Preview for nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes
2026

nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes

PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card Dataset Description PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.

license:othersize_categories:100K<n<1Mmodality:videoregion:usroboticssynthetic-data
Preview for nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes
2026

nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes

PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card Dataset Description PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding. Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.

license:othersize_categories:100M<n<1Bformat:webdatasetmodality:imagemodality:textlibrary:datasets
Preview for nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes
2026

nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes

PhysicalAI SDG-Warehouse PhysicalAI SDG-Warehouse is a synthetic, fully-annotated video dataset of staged industrial-safety events captured in a simulated warehouse environment. It contains approximately 123k video clips, totaling roughly 412 hours of footage at 1920x1080 resolution and 30 frames per second, organized across four scenarios: a forklift near-miss with a human worker, a warehouse fire with worker evacuation, a forklift collision with a storage shelf, and a routine… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes.

task_categories:video-classificationtask_categories:video-text-to-texttask_categories:text-to-videolanguage:enlicense:othersize_categories:100K<n<1M
Preview for open-gigaai/CVPR-2026-WorldModel-Track-Dataset
2026

open-gigaai/CVPR-2026-WorldModel-Track-Dataset

GigaBrain Challenge 2026 (CVPR 2026 Workshop Competition) Registration To access the dataset you must register your team. Required information: Team name Team leader Team members Organization Leader email Click Request Access to participate. Resources After approval you will be able to download: Training dataset Test dataset Baseline model Evaluation scripts

region:us
Preview for PatronusAI/world_model_corpus
2026

PatronusAI/world_model_corpus

Dataset Card for World Model Corpus The world model corpus contains a set of generated trajectories that are shaped for text-based world modeling task as used by the paper: "Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL". The dataset contains trajectories from nine distinct environments: Tau2Bench, SWE-Smith, DeepresearchQA, Openresearcher, Gorilla/BFCLv4, Webshop, Toolathlon, Pandora and Coderforge. Loading from… See the full description on the dataset page: https://huggingface.co/datasets/PatronusAI/world_model_corpus.

task_categories:text-generationlicense:mitsize_categories:100K<n<1Mformat:parquetmodality:textlibrary:datasets
Preview for shubhxho/ego-world-model-v1
2026

shubhxho/ego-world-model-v1

language: en license: cc-by-4.0 size_categories: 10K<n<100K source_datasets: MicroAGI-Labs/MicroAGI00 task_categories: video-classification image-segmentation depth-estimation object-detection pretty_name: Ego World Model Dataset v1 tags: ego-centric world-model robotics sam3 depth-estimation optical-flow action-conditioned first-person-video--- Ego World Model Dataset v1 Egocentric RGB + metric depth + SAM 3 instance segmentation + action labels, built for training… See the full description on the dataset page: https://huggingface.co/datasets/shubhxho/ego-world-model-v1.

region:us
Preview for twelvedata/financial-world-model
2026

twelvedata/financial-world-model

Twelve Data World Model Dataset A multi-modal financial time-series dataset built from Twelve Data market data. Each timeframe is published in three parallel views: bars_* — OHLCV bars enriched with causal technical indicators and macro context, in Parquet. text_* — instruction-tuning prompts/labels derived from the bars, in JSONL. trajectories_* — fixed-length rolling windows of state vectors plus next-state pairs, suitable for world-model / sequence-model training, in… See the full description on the dataset page: https://huggingface.co/datasets/twelvedata/financial-world-model.

task_categories:time-series-forecastingtask_categories:text-generationtask_categories:reinforcement-learninglanguage:enlicense:mitsize_categories:10M<n<100M
Preview for VideoWorldmodel/ReasoningStructureTestset
2026

VideoWorldmodel/ReasoningStructureTestset

Reasoning-Structured Videos A Stratified Diagnostic Suite for Compositional Consistency in Action-Conditioned Video World Models. Reasoning-Structured Videos is a UE5-rendered video benchmark whose trajectories are organised as rooted graphs with path-level algebraic relations. Unlike flat corpora that release independent action–observation rollouts, every released trajectory here is annotated as an exact instance of one of three identities a faithful transition operator must… See the full description on the dataset page: https://huggingface.co/datasets/VideoWorldmodel/ReasoningStructureTestset.

task_categories:video-classificationtask_categories:otherlanguage:enlicense:cc-by-4.0size_categories:1K<n<10Kregion:us
Preview for Wh0
2026

Wh0

Wh0 uses generative video world models to produce scalable egocentric human-hand manipulation data, improving dexterous VLA model zero-shot success on real-world tasks.

generative world modelsegocentric videodexterous manipulationdata generationhuman-object interactionVLA models
Preview for xwk123/Mobile-GUI-Worldmodel-SFT
2026

xwk123/Mobile-GUI-Worldmodel-SFT

Mobile-GUI-Worldmodel-SFT This repository contains mobile GUI agent data and auxiliary files for training and evaluating GUI world models. The data is organized around GUI trajectories: each step has a screenshot and page-state annotations such as HTML, plain text, and structured text. Repository Layout . ├── GUI-agent-main/ # Data annotation scripts and examples ├── eval/ # Evaluation assets │ └── AndroidControl_images.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/xwk123/Mobile-GUI-Worldmodel-SFT.

task_categories:image-to-texttask_categories:visual-question-answeringtask_categories:roboticslanguage:enlicense:othersize_categories:10K<n<100K
Preview for ywang077/Trajectory_world_model_dataset
2026

ywang077/Trajectory_world_model_dataset

WestWorld Pretraining Dataset This repository contains the pretraining dataset for WestWorld, a knowledge-encoded scalable trajectory world model for diverse robotic systems. Paper | Project Page | GitHub Description WestWorld is designed to address the scalability challenges in trajectory world models for diverse robotic systems. The dataset includes trajectories from 89 complex environments spanning diverse morphologies across both simulation and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ywang077/Trajectory_world_model_dataset.

task_categories:roboticslicense:cc-by-nc-4.0arxiv:2603.14392region:us
Preview for ZaidGhazal/world-models-eval
2026

ZaidGhazal/world-models-eval

DreamGrasp: Processed LIBERO Manipulation Demonstrations Does a robot policy's evaluation still mean something if it never touched a real simulator, only a world model's imagination of one? This dataset is the shared training data behind that question, a single, ready-to-train release built from LIBERO's manipulation demonstrations (libero_spatial, libero_object, libero_goal). It provides: Fixed, versioned train / validation / test / held-out splits, so every result trained on… See the full description on the dataset page: https://huggingface.co/datasets/ZaidGhazal/world-models-eval.

task_categories:roboticslicense:mitsize_categories:100K<n<1Mformat:parquetmodality:tabularmodality:timeseries
Preview for 1x-technologies/world_model_raw_data
2025

1x-technologies/world_model_raw_data

Raw Dataset for the 1X World Model Sammpling Challenge. Download with: huggingface-cli download 1x-technologies/worldmodel_raw_data --repo-type dataset --local-dir data Train/Val v2.0 The training dataset is shareded into 100 independent shards. The definitions are as follows: video_{shard}.mp4: Raw video with a resolution of 512x512. segment_idx_{shard}.bin - Maps each frame i to its corresponding segment index. You may want to use this to separate non-contiguous frames from… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_raw_data.

license:cc-by-nc-sa-4.0size_categories:10M<n<100Mregion:us
Preview for 1x-technologies/world_model_tokenized_data
2025

1x-technologies/world_model_tokenized_data

1X World Model Compression Challenge Dataset This repository hosts the dataset for the 1X World Model Compression Challenge. huggingface-cli download 1x-technologies/worldmodel --repo-type dataset --local-dir data Updates Since v1.1 Train/Val v2.0 (~100 hours), replacing v1.1 Test v2.0 dataset for the Compression Challenge Faces blurred for privacy New raw video dataset (CC-BY-NC-SA 4.0) at worldmodel_raw_data Example scripts now split into: cosmos_video_decoder.py —… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_tokenized_data.

license:apache-2.0size_categories:10M<n<100Mregion:us
Preview for Clementppr/lerobot_pick_and_place_dataset_world_model
2025

Clementppr/lerobot_pick_and_place_dataset_world_model

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 30, "total_frames": 13572, "total_tasks": 1, "total_videos": 30, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:30"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Clementppr/lerobot_pick_and_place_dataset_world_model.

task_categories:roboticslicense:apache-2.0size_categories:10K<n<100Kformat:parquetmodality:tabularmodality:timeseries
Preview for GEM
2025

GEM

GEM is a multimodal world model for controllable future frame prediction with ego-motion, object dynamics, and human pose control, using a large real-world dataset.

Preview for ORV
2025

ORV

ORV introduces a 4D occupancy-centric framework for controllable robot video generation, improving fidelity, temporal consistency, and control alignment.

Preview for thuml/bytesized32-world-model-cot
2025

thuml/bytesized32-world-model-cot

See https://github.com/thuml/RLVR-World for examples for using this dataset. Citation @article{wu2025rlvr, title={RLVR-World: Training World Models with Reinforcement Learning}, author={Jialong Wu and Shaofeng Yin and Ningya Feng and Mingsheng Long}, journal={arXiv preprint arXiv:2505.13934}, year={2025}, }

license:mitsize_categories:100K<n<1Mformat:parquetmodality:textlibrary:datasetslibrary:dask
Preview for EVA
2024

EVA

Proposes EVA, an embodied world model using vision-language and video generation models for multi-step video prediction and OOD handling.