
数据集大全
数据集
按表示方法和任务分组的数据集。




迈向交互式视频世界建模
交互式视频世界建模的综述,涵盖前沿、挑战、基准和未来趋势。





用于自动驾驶场景预测的扩散Transformer世界-动作模型
一种紧凑的潜在扩散Transformer世界模型,用于自动驾驶场景预测,以自车动作为条件,在感知指标上优于回归方法。





Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Testl
该论文围绕“Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Testl”研究世界模型相关问题。




Physics-IQ 验证
审计并改进Physics-IQ基准,用于评估视频生成模型的物理理解,优化提示和评分。



Qwen-RobotWorld技术报告
Qwen-RobotWorld是语言条件视频世界模型,统一多个具身领域,在多项基准测试中排名第一。

stable-worldmodel
一个用于可复现世界模型研究的开源平台,提供标准化数据、基线和评估基准。



重建还是语义?什么使潜在空间对机器人世界模型有用
比较重建与语义潜在空间在动作条件视频扩散世界模型中的应用,发现语义编码器对策略相关机器人更优。


WorldOlympiad
WorldOlympiad基准测试从物理、几何和交互三个维度评估视频世界模型,覆盖游戏、机器人和真实场景。


WorldReasonBench
提出WorldReasonBench基准,测试视频生成器作为世界状态预测器,揭示视觉合理性与推理差距。



Holo-World
Holo-World从单张图像统一控制相机、物体和天气,通过数据集和新颖适配器生成视频世界模型。


当前世界模型缺乏持久状态核心
本文提出WRBench基准,测试世界模型在未被观测时是否维持持久状态,发现当前模型在遮挡期间无法推进事件演化。





机器人操作的世界模型
综述机器人操作的世界模型,将其定义为动作条件预测系统,并按表示、预测-动作耦合和使用流程组织方法。



你的驾驶世界模型是全能选手吗?
提出WorldLens统一基准,从像素质量、几何、闭环规划及人类感知评估驾驶世界模型。







OccDirector
OccDirector 从自然语言生成4D占用动态,实现最先进的指令跟随。

τ₀-WM
一种统一视频-动作世界模型,集成策略学习、视频预测与动作评估,用于机器人操作。


Prisma-World
Prisma-World是一个可通过相机控制的多智能体视频世界模型,通过几何感知去噪和新合成数据集确保跨视角一致性。







$\mu_0$
μ0从视频预测3D交互轨迹,无需动作标签实现可扩展机器人学习。





修剪视觉世界建模评估的长尾
引入Tailor-Bench评估视觉世界模型在长尾物理交互上的表现,揭示其在常见场景之外的泛化能力有限。


PhyGround
PhyGround通过250个提示、13条物理定律和大规模人类研究,对生成式世界模型中的物理推理进行基准测试。


针对World Models以破坏机器人学习流水线
世界模型引入隐蔽数据投毒,使安全训练数据仍导致策略被破坏。




ReactSim-Bench
提出ReactSim-Bench,用解耦控制和多种指标评估自动驾驶行为世界模型反应能力。



基于世界模型的视频生成中物理一致性的无参考评估
提出基于DROID-SLAM和SEA-RAFT的无参考物理一致性评估方法,提升视频生成任务成功率8%。




NarrativeWorldBench
提出长程音频剧基准NarrativeWorldBench和潜世界模型N-VSSM,后者在200集内保持高一致性,超越前沿LLM。

EgoCS-400K
提出EgoCS-400K,大规模第一人称游戏数据集,含对齐视频-动作-语言轨迹,用于训练交互世界模型。






ReflectiChain
ReflectiChain通过生成式供应链世界模型和双环学习弥合LLM与RL差距,在半导体基准上提升韧性。


将游戏代码世界模型生成蒸馏到轻量级大语言模型中
通过SFT和RLVR将游戏代码世界模型生成能力蒸馏到小型LLM中,提升语法正确性和规则遵循度。

LoViF 2026 首届面向4D世界模型的整体质量评估挑战赛 (PhyScore)
LoViF 2026 PhyScore挑战赛对世界模型生成的视频进行整体质量评估,评估物理真实性、时间一致性和异常定位。

面向几何一致性的定量视频世界模型评估
提出PDI-Bench,一种定量评估生成视频几何一致性的框架,揭示感知指标无法捕捉的失败模式。





WorldArena 2.0
WorldArena 2.0 在模态、功能和平台三个维度扩展了具身世界模型基准测试,并包含真实世界评估。





移动世界模型如何指导GUI智能体?
研究移动世界模型如何通过四种模态引导GUI代理,并评估其下游效用。











利用真实世界视频学习粒子动力学模型
利用高斯泼溅和渲染监督从真实世界视频学习粒子动力学,克服模拟到真实差距。


通过对话对齐世界模型的具身多智能体协调
基于LLM的具身智能体通过对话对齐世界模型,减少冲突但降低任务成功率;新指标衡量对齐差距。
























面向人形机器人的生成式世界建模
该论文围绕“Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report”研究世界模型相关问题。


















TC-Bench
Derived from paper: TC-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation








WorldRoamBench
尽管交互式世界模型(IWM)取得了快速进展,现有基准仅在轨迹层面评估动作跟随,忽略了记忆与交互物理。我们提出WorldRoamBench,一个面向长时程稳定性的开放世界基准,涵盖四个维度,每个维度均有定制创新:(i)动作:逐帧动作度量,绕过跨模型语义尺度差异,暴露被轨迹隐藏的失败;(ii)视觉:基于片段的漂移度量,捕捉起始-终点比较无法发现的非单调中期序列崩溃;(iii)物理:可控性门控评估,涵盖力学、光学与3D一致性,在忠实动作执行下对合理性进行评分;(iv)记忆:动作解耦协议,通过过渡定位的3D点云重建评估场景记忆,通过跟踪加VLM推理评估主体记忆。该基准包含600余个测试用例,覆盖自然、城市与室内场景,采用第一/第三人称视角,支持WASD连续交互10-60秒。对10余个开源/闭源模型的评估显示,没有任何模型能可靠满足所有维度;即使最佳模型也仅获得中等分数。在WorldRoamBench上的进步是迈向稳定、物理可信、记忆忠实且可部署于实际应用的交互式世界模型的重要步骤。


DreamTraj
在操作过程中准确预测物体轨迹对于闭环感知-动作循环至关重要。当前进展受限于两个方面:现有数据集缺乏细粒度的语言-运动标注,且现有预测器要么依赖视频、深度或CAD模型等特权输入,要么通过昂贵且易出错的感知管道从完全生成的视频中恢复运动。我们通过MOVE数据集弥补了监督差距,该数据集包含5,038条以物体为中心的自我中心轨迹,每条轨迹均配有细粒度的自然语言指令,而非粗略的动词-名词标签。我们进一步提出了DreamTraj,它仅从单张RGB图像和任务指令即可预测6自由度物体轨迹,推理时无需视频、深度或CAD模型:DreamTraj不生成视频,而是从冻结的图像到视频扩散模型在早期去噪步骤的内部表示中读取运动。一个轻量级的流匹配解码器(Reader)将查询-键注意力轨迹和池化隐藏状态解码为相对6自由度位姿。据我们所知,这是首个直接从中间视频扩散表示而非生成像素解码物体6自由度轨迹的方法。DreamTraj在平移和旋转方面均达到了新的最先进水平,超越了使用多帧或特权输入的预测器,并且运行速度比先生成再提取的管道快4.6倍。




RoboTrustBench
Derived from paper: RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
ChangChrisLiu/GNN_Disassembly_WorldModel
GNN Constraint-Aware World Model Dataset (v3) Real robot episodes with per-frame constraint graphs, SAM2 segmentation masks + 256-D feature embeddings, full 3D depth bundles, and synchronized robot states across two manipulation domains. Both domains share the v3 on-disk layout (same JSON/NPZ schemas, same delta-encoded frame_states, same fully-connected PyG expansion at load time) and now share a unified 270-D node feature format — the PyG loader reads a fixed 10-D type… See the full description on the dataset page: https://huggingface.co/datasets/ChangChrisLiu/GNN_Disassembly_WorldModel.
GroupToM-Bench
Introduces GroupToM-Bench, a multimodal benchmark evaluating group-level theory of mind in MLLMs, revealing gaps in social world modeling.
Kalso42/WorldModelForMaze
WorldModelForMaze Code, datasets, and trained checkpoints for studying world-model representations in maze navigation, based on a modified NanoGPT. Contents *.py — training, testing, probing, and visualization scripts (see readme.md). model/ — architectures: transformer, transformer-rope, transformer-nextlat, mamba, mamba2, gated-deltanet, gru. data/maze/100/ — tokenized maze datasets for Tasks A/C/E/H/I (RWs paths, 100 nodes). out/ — final (10000-iter)… See the full description on the dataset page: https://huggingface.co/datasets/Kalso42/WorldModelForMaze.
Kalso42/WorldModelForMazeWithX
WorldModelForMazeWithX Maze pathfinding sequences for training/probing sequence models (Transformer, Mamba, GRU, Gated-DeltaNet, ...). Task C1: relative-turn navigation on a fixed 10×10 directed grid. Includes a special x terminator marking wall-hit (illegal) paths, used to study a model's ability to recognize its own errors. Maze 10×10 grid, 100 nodes (0–99). Directed edges (down/right, both directions added), edge probability 0.6. Graph:… See the full description on the dataset page: https://huggingface.co/datasets/Kalso42/WorldModelForMazeWithX.
kfallah/world-model-pi-sft-with-cot
world-model-pi-sft-with-cot SFT-ready data for training a world model of the TeichAI pi coding-agent harness. Each world-model assistant target carries a teacher-distilled <think>{rationale}</think> block explaining why the next environment turn follows from prior context, followed by the original [developer] / [tool:*] environment block. Stage 2 of a 2-stage pipeline. Stage 1 dataset (with <COT_PLACEHOLDER> slots) is at kfallah/world-model-pi-sft-formatted. Schema… See the full description on the dataset page: https://huggingface.co/datasets/kfallah/world-model-pi-sft-with-cot.
nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios
Dataset Description: PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios is a large-scale synthetic video dataset of autonomous-driving scenes generated with NVIDIA's internal Omniverse simulation platform. Each clip is a temporally consistent multi-camera surround capture of one ego vehicle and surrounding traffic participants, paired with per-camera VLM captions. The dataset is designed to fill gaps in real-world driving data along two axes: (1) targeted long-tail… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios.
nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes
Dataset Description: The SDG-SynHuman is a large-scale synthetic video dataset of digital humans rendered in diverse indoor and outdoor 3D environments. The dataset contains 236,937 clips, totaling approximately 5,841 hours of video, and is designed to support training and post-training of NVIDIA Cosmos world foundation models and related physical AI research. Each sample is a temporally coherent 60-120 second video clip rendered at 1080p and 30 fps. Clips contain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes.
nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes
PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card Dataset Description PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.
nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card Dataset Description PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding. Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.
nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes
PhysicalAI SDG-Warehouse PhysicalAI SDG-Warehouse is a synthetic, fully-annotated video dataset of staged industrial-safety events captured in a simulated warehouse environment. It contains approximately 123k video clips, totaling roughly 412 hours of footage at 1920x1080 resolution and 30 frames per second, organized across four scenarios: a forklift near-miss with a human worker, a warehouse fire with worker evacuation, a forklift collision with a storage shelf, and a routine… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes.
open-gigaai/CVPR-2026-WorldModel-Track-Dataset
GigaBrain Challenge 2026 (CVPR 2026 Workshop Competition) Registration To access the dataset you must register your team. Required information: Team name Team leader Team members Organization Leader email Click Request Access to participate. Resources After approval you will be able to download: Training dataset Test dataset Baseline model Evaluation scripts
PatronusAI/world_model_corpus
Dataset Card for World Model Corpus The world model corpus contains a set of generated trajectories that are shaped for text-based world modeling task as used by the paper: "Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL". The dataset contains trajectories from nine distinct environments: Tau2Bench, SWE-Smith, DeepresearchQA, Openresearcher, Gorilla/BFCLv4, Webshop, Toolathlon, Pandora and Coderforge. Loading from… See the full description on the dataset page: https://huggingface.co/datasets/PatronusAI/world_model_corpus.
shubhxho/ego-world-model-v1
language: en license: cc-by-4.0 size_categories: 10K<n<100K source_datasets: MicroAGI-Labs/MicroAGI00 task_categories: video-classification image-segmentation depth-estimation object-detection pretty_name: Ego World Model Dataset v1 tags: ego-centric world-model robotics sam3 depth-estimation optical-flow action-conditioned first-person-video--- Ego World Model Dataset v1 Egocentric RGB + metric depth + SAM 3 instance segmentation + action labels, built for training… See the full description on the dataset page: https://huggingface.co/datasets/shubhxho/ego-world-model-v1.
syCen/action-worldmodel-bench
Hugging Face dataset: syCen/action-worldmodel-bench
twelvedata/financial-world-model
Twelve Data World Model Dataset A multi-modal financial time-series dataset built from Twelve Data market data. Each timeframe is published in three parallel views: bars_* — OHLCV bars enriched with causal technical indicators and macro context, in Parquet. text_* — instruction-tuning prompts/labels derived from the bars, in JSONL. trajectories_* — fixed-length rolling windows of state vectors plus next-state pairs, suitable for world-model / sequence-model training, in… See the full description on the dataset page: https://huggingface.co/datasets/twelvedata/financial-world-model.
VideoWorldmodel/ReasoningStructureTestset
Reasoning-Structured Videos A Stratified Diagnostic Suite for Compositional Consistency in Action-Conditioned Video World Models. Reasoning-Structured Videos is a UE5-rendered video benchmark whose trajectories are organised as rooted graphs with path-level algebraic relations. Unlike flat corpora that release independent action–observation rollouts, every released trajectory here is annotated as an exact instance of one of three identities a faithful transition operator must… See the full description on the dataset page: https://huggingface.co/datasets/VideoWorldmodel/ReasoningStructureTestset.
VideoWorldmodel/ReasoningVideoSamples
Hugging Face dataset: VideoWorldmodel/ReasoningVideoSamples

xwk123/Mobile-GUI-Worldmodel-SFT
Mobile-GUI-Worldmodel-SFT This repository contains mobile GUI agent data and auxiliary files for training and evaluating GUI world models. The data is organized around GUI trajectories: each step has a screenshot and page-state annotations such as HTML, plain text, and structured text. Repository Layout . ├── GUI-agent-main/ # Data annotation scripts and examples ├── eval/ # Evaluation assets │ └── AndroidControl_images.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/xwk123/Mobile-GUI-Worldmodel-SFT.
ywang077/Trajectory_world_model_dataset
WestWorld Pretraining Dataset This repository contains the pretraining dataset for WestWorld, a knowledge-encoded scalable trajectory world model for diverse robotic systems. Paper | Project Page | GitHub Description WestWorld is designed to address the scalability challenges in trajectory world models for diverse robotic systems. The dataset includes trajectories from 89 complex environments spanning diverse morphologies across both simulation and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ywang077/Trajectory_world_model_dataset.
ZaidGhazal/world-models-eval
DreamGrasp: Processed LIBERO Manipulation Demonstrations Does a robot policy's evaluation still mean something if it never touched a real simulator, only a world model's imagination of one? This dataset is the shared training data behind that question, a single, ready-to-train release built from LIBERO's manipulation demonstrations (libero_spatial, libero_object, libero_goal). It provides: Fixed, versioned train / validation / test / held-out splits, so every result trained on… See the full description on the dataset page: https://huggingface.co/datasets/ZaidGhazal/world-models-eval.
1x-technologies/world_model_raw_data
Raw Dataset for the 1X World Model Sammpling Challenge. Download with: huggingface-cli download 1x-technologies/worldmodel_raw_data --repo-type dataset --local-dir data Train/Val v2.0 The training dataset is shareded into 100 independent shards. The definitions are as follows: video_{shard}.mp4: Raw video with a resolution of 512x512. segment_idx_{shard}.bin - Maps each frame i to its corresponding segment index. You may want to use this to separate non-contiguous frames from… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_raw_data.
1x-technologies/world_model_tokenized_data
1X World Model Compression Challenge Dataset This repository hosts the dataset for the 1X World Model Compression Challenge. huggingface-cli download 1x-technologies/worldmodel --repo-type dataset --local-dir data Updates Since v1.1 Train/Val v2.0 (~100 hours), replacing v1.1 Test v2.0 dataset for the Compression Challenge Faces blurred for privacy New raw video dataset (CC-BY-NC-SA 4.0) at worldmodel_raw_data Example scripts now split into: cosmos_video_decoder.py —… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_tokenized_data.
Clementppr/lerobot_pick_and_place_dataset_world_model
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 30, "total_frames": 13572, "total_tasks": 1, "total_videos": 30, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:30"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Clementppr/lerobot_pick_and_place_dataset_world_model.


thuml/bytesized32-world-model-cot
See https://github.com/thuml/RLVR-World for examples for using this dataset. Citation @article{wu2025rlvr, title={RLVR-World: Training World Models with Reinforcement Learning}, author={Jialong Wu and Shaofeng Yin and Ningya Feng and Mingsheng Long}, journal={arXiv preprint arXiv:2505.13934}, year={2025}, }
