
数据集大全
数据集
按表示方法和任务分组的数据集。




迈向交互式视频世界建模
交互式视频世界建模的综述,涵盖前沿、挑战、基准和未来趋势。





用于自动驾驶场景预测的扩散Transformer世界-动作模型
一种紧凑的潜在扩散Transformer世界模型,用于自动驾驶场景预测,以自车动作为条件,在感知指标上优于回归方法。





Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Testl
该论文围绕“Wow, wo, val! A Comprehensive Embodied World Model Evaluation Turing Testl”研究世界模型相关问题。




Physics-IQ 验证
审计并改进Physics-IQ基准,用于评估视频生成模型的物理理解,优化提示和评分。



Qwen-RobotWorld技术报告
Qwen-RobotWorld是语言条件视频世界模型,统一多个具身领域,在多项基准测试中排名第一。

stable-worldmodel
一个用于可复现世界模型研究的开源平台,提供标准化数据、基线和评估基准。



重建还是语义?什么使潜在空间对机器人世界模型有用
比较重建与语义潜在空间在动作条件视频扩散世界模型中的应用,发现语义编码器对策略相关机器人更优。


WorldOlympiad
WorldOlympiad基准测试从物理、几何和交互三个维度评估视频世界模型,覆盖游戏、机器人和真实场景。


WorldReasonBench
提出WorldReasonBench基准,测试视频生成器作为世界状态预测器,揭示视觉合理性与推理差距。



Holo-World
Holo-World从单张图像统一控制相机、物体和天气,通过数据集和新颖适配器生成视频世界模型。


当前世界模型缺乏持久状态核心
本文提出WRBench基准,测试世界模型在未被观测时是否维持持久状态,发现当前模型在遮挡期间无法推进事件演化。





机器人操作的世界模型
综述机器人操作的世界模型,将其定义为动作条件预测系统,并按表示、预测-动作耦合和使用流程组织方法。



你的驾驶世界模型是全能选手吗?
提出WorldLens统一基准,从像素质量、几何、闭环规划及人类感知评估驾驶世界模型。







OccDirector
OccDirector 从自然语言生成4D占用动态,实现最先进的指令跟随。

τ₀-WM
一种统一视频-动作世界模型,集成策略学习、视频预测与动作评估,用于机器人操作。


Prisma-World
Prisma-World是一个可通过相机控制的多智能体视频世界模型,通过几何感知去噪和新合成数据集确保跨视角一致性。







$\mu_0$
μ0从视频预测3D交互轨迹,无需动作标签实现可扩展机器人学习。





修剪视觉世界建模评估的长尾
引入Tailor-Bench评估视觉世界模型在长尾物理交互上的表现,揭示其在常见场景之外的泛化能力有限。


PhyGround
PhyGround通过250个提示、13条物理定律和大规模人类研究,对生成式世界模型中的物理推理进行基准测试。


针对World Models以破坏机器人学习流水线
世界模型引入隐蔽数据投毒,使安全训练数据仍导致策略被破坏。




ReactSim-Bench
提出ReactSim-Bench,用解耦控制和多种指标评估自动驾驶行为世界模型反应能力。



基于世界模型的视频生成中物理一致性的无参考评估
提出基于DROID-SLAM和SEA-RAFT的无参考物理一致性评估方法,提升视频生成任务成功率8%。




NarrativeWorldBench
提出长程音频剧基准NarrativeWorldBench和潜世界模型N-VSSM,后者在200集内保持高一致性,超越前沿LLM。

EgoCS-400K
提出EgoCS-400K,大规模第一人称游戏数据集,含对齐视频-动作-语言轨迹,用于训练交互世界模型。






ReflectiChain
ReflectiChain通过生成式供应链世界模型和双环学习弥合LLM与RL差距,在半导体基准上提升韧性。

将游戏代码世界模型生成蒸馏到轻量级大语言模型中
通过SFT和RLVR将游戏代码世界模型生成能力蒸馏到小型LLM中,提升语法正确性和规则遵循度。

LoViF 2026 首届面向4D世界模型的整体质量评估挑战赛 (PhyScore)
LoViF 2026 PhyScore挑战赛对世界模型生成的视频进行整体质量评估,评估物理真实性、时间一致性和异常定位。

面向几何一致性的定量视频世界模型评估
提出PDI-Bench,一种定量评估生成视频几何一致性的框架,揭示感知指标无法捕捉的失败模式。





WorldArena 2.0
WorldArena 2.0 在模态、功能和平台三个维度扩展了具身世界模型基准测试,并包含真实世界评估。





移动世界模型如何指导GUI智能体?
研究移动世界模型如何通过四种模态引导GUI代理,并评估其下游效用。












利用真实世界视频学习粒子动力学模型
利用高斯泼溅和渲染监督从真实世界视频学习粒子动力学,克服模拟到真实差距。


通过对话对齐世界模型的具身多智能体协调
基于LLM的具身智能体通过对话对齐世界模型,减少冲突但降低任务成功率;新指标衡量对齐差距。






















PAWBench
近期的视频生成模型越来越多地被视作世界模型。许多物理过程可以以不止一种有效方式展开。因此,世界模型不仅应重现一条合理的轨迹,还应重现相同初始观测和动作下可能行为的分布。我们将这种分布层面的要求称为概率对齐(probabilistic alignment)。然而,现有评估大多只考察单个视频的合理性,并未检验重复生成是否能够恢复正确的分布。这引出了一个核心问题:当前的视频生成器距离概率对齐的世界建模还有多远?为了回答这一问题,我们将概率对齐形式化为世界模型的一个分布性准则,并引入 PAWBench——一个将视频生成器作为世界动力学随机采样器进行评估的基准。我们进一步提出 PAWEval,一种结果层面的评估协议,将重复的视频推演转换为可能物理行为上的经验分布。在50个场景和11个当前系统中,没有任何模型能够在恢复有效行为范围的同时,始终匹配参考概率。在确认这一差距后,我们检验了语言提示、初始噪声采样或模型训练是否能重塑模型的预测分布。我们相信,这项工作可以为未来迈向概率对齐世界建模的努力奠定基础。






面向人形机器人的生成式世界建模
该论文围绕“Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report”研究世界模型相关问题。





WorldRover
学习生成或重建可探索世界,需要视频不仅配准RGB,还需要相机运动、场景几何、时间对应关系,以及对于交互式模型而言的控制信号。真实采集可以提供其中部分信号,但稠密几何和长程对应通常依赖估计或专门仪器。渲染可以直接提供这些量,但现有的合成资源很少在同一帧上同时包含它们,同时支持对视角和外观的受控变化。我们提出WorldRover,一个用于生成艺术家构建环境的丰富标注、长程探索的数据引擎。其核心WorldRover-Engine是一个Unreal Engine流水线,能够执行并离线渲染分钟级路线,同时保留完整的轨迹和场景几何。同一探索可以在不同环境状态下,从第一人称、第三人称和360度全景相机重放。利用WorldRover-Engine,我们构建了WorldRover-10M,其序列在每次探索中将RGB与度量深度、相机轨迹以及由轨迹导出的动作信号配对。第三人称子集还额外提供稠密光流、带可见性的长程2D/3D点轨迹,以及与相机轨迹不同的角色轨迹。该引擎可以在不同环境状态下或使用中性白色材质,从第一人称、第三人称和360度全景视角渲染一次遍历,同时保留路线和场景几何。因此,WorldRover将长视界世界探索转化为一个可扩展的数据生成问题,为必须构建、维护并重新访问可探索世界的一致表征的模型提供监督。














CineScene
电影级视频制作需要控制场景-主体构图和摄像机运动,但实景拍摄由于需要搭建实体布景而成本高昂。为解决这一问题,我们提出了场景上下文解耦的电影级视频生成任务:给定静态环境的多个图像,目标是合成以动态主体为特征的高质量视频,同时保持底层场景一致性并遵循用户指定的摄像机轨迹。我们提出了CineScene,一个利用隐式3D感知场景表示进行电影级视频生成的框架。我们的关键创新是一种新颖的上下文条件机制,以隐式方式注入3D感知特征:通过VGGT将场景图像编码为视觉表示,CineScene通过额外的上下文拼接将空间先验注入预训练的文本到视频生成模型,从而实现在一致场景和动态主体下的摄像机可控视频合成。为了进一步增强模型鲁棒性,我们引入了一种简单而有效的训练时输入场景图像随机洗牌策略。为解决训练数据缺乏的问题,我们使用Unreal Engine 5构建了一个场景解耦数据集,包含有无动态主体的场景配对视频、代表底层静态场景的全景图像及其摄像机轨迹。实验表明,CineScene在场景一致的电影级视频生成方面达到了最先进性能,能够处理大范围摄像机运动,并展示了对多样环境的泛化能力。





TC-Bench
Derived from paper: TC-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation











4DAnyone
我们提出了4DAnyone,一个从未标定的单目视频中重建4D人体的框架,其通过生成重建级的多视角一致视频,并将其提升为4D高斯泼溅(4DGS)表示。现有的相机控制视频扩散模型能够合成看似合理的新视角视频,但在扩展到4DGS重建所需的数十个目标视角时,无法保持一致性。我们将这一失败归因于有界注意力上下文问题:当目标视角超过单次DiT前向传播的容量时,必须将其分成若干组,从而暴露出两个相互耦合的瓶颈。在参考上下文方面,对所有先前生成视角的条件依赖以O(N)的速度增长,削弱了跨视角外观引导。在目标上下文方面,互不相交的组无法直接交换信息,导致全局结构漂移。4DAnyone通过两种互补设计解决这两个瓶颈:参考上下文打包(RCP)将不断增长的参考视角压缩为固定长度的混合分辨率上下文,实现O(1)的参考上下文复杂度;目标上下文路由(TCR)在去噪过程中轮换目标视角的分组,从而在高噪声步骤中跨组共享上下文,并在低噪声步骤中稳定细节。我们进一步利用自研游戏引擎构建了MVGameHuman数据集,并将其与light-stage和in-the-wild视频数据集结合用于训练。在DNA-Rendering和DyMVHumans上的实验表明,4DAnyone在新视角视频质量和下游4DGS重建方面均优于先前方法,并具有强大的in-the-wild泛化能力。更多视频结果和源代码请参见我们的项目页面:https://4danyone.github.io。




WorldRoamBench
尽管交互式世界模型(IWM)取得了快速进展,现有基准仅在轨迹层面评估动作跟随,忽略了记忆与交互物理。我们提出WorldRoamBench,一个面向长时程稳定性的开放世界基准,涵盖四个维度,每个维度均有定制创新:(i)动作:逐帧动作度量,绕过跨模型语义尺度差异,暴露被轨迹隐藏的失败;(ii)视觉:基于片段的漂移度量,捕捉起始-终点比较无法发现的非单调中期序列崩溃;(iii)物理:可控性门控评估,涵盖力学、光学与3D一致性,在忠实动作执行下对合理性进行评分;(iv)记忆:动作解耦协议,通过过渡定位的3D点云重建评估场景记忆,通过跟踪加VLM推理评估主体记忆。该基准包含600余个测试用例,覆盖自然、城市与室内场景,采用第一/第三人称视角,支持WASD连续交互10-60秒。对10余个开源/闭源模型的评估显示,没有任何模型能可靠满足所有维度;即使最佳模型也仅获得中等分数。在WorldRoamBench上的进步是迈向稳定、物理可信、记忆忠实且可部署于实际应用的交互式世界模型的重要步骤。


H2R-Bench
大规模操作数据对于机器人学习至关重要,然而收集机器人示范数据仍然昂贵且难以扩展。与此同时,大量第一人称人类操作视频提供了丰富的行为经验,但由于人类手部与机器人末端执行器之间的差异,跨具身迁移仍然具有挑战性。视频世界模型的最新进展为从人类观察中合成以机器人为中心的操作视频提供了一条有前景的路径,但其跨具身迁移能力在很大程度上尚未被探索。因此,我们提出了 H2R-Bench,一个用于评估跨具身人类到机器人操作视频生成的基准,其中模型在指定具身条件下将第一人称人类示范转换为机器人操作视频。每个基准实例包含一段人类示范视频、目标具身约束,以及覆盖任务目标、动作事件、功能接触和物体响应的基于源视频的标注。H2R-Bench 通过五个维度评估生成的视频,包括目标状态完成度、动作事件完成度、功能接触迁移、具身正确性和通用视频质量。我们对十一个最先进的视频生成模型在六个操作任务族和两种机器人具身上进行了基准测试。我们的评估表明,当前的视频世界模型在人类到机器人操作迁移方面仍然有限:即使是领先模型也常在具身一致性、功能交互和任务执行上失败。H2R-Bench 提供了一个系统化的诊断框架,用于评估视频世界模型能否弥合人类到机器人的具身差距,并将人类操作观察转化为以机器人为中心的训练资源。



DreamTraj
在操作过程中准确预测物体轨迹对于闭环感知-动作循环至关重要。当前进展受限于两个方面:现有数据集缺乏细粒度的语言-运动标注,且现有预测器要么依赖视频、深度或CAD模型等特权输入,要么通过昂贵且易出错的感知管道从完全生成的视频中恢复运动。我们通过MOVE数据集弥补了监督差距,该数据集包含5,038条以物体为中心的自我中心轨迹,每条轨迹均配有细粒度的自然语言指令,而非粗略的动词-名词标签。我们进一步提出了DreamTraj,它仅从单张RGB图像和任务指令即可预测6自由度物体轨迹,推理时无需视频、深度或CAD模型:DreamTraj不生成视频,而是从冻结的图像到视频扩散模型在早期去噪步骤的内部表示中读取运动。一个轻量级的流匹配解码器(Reader)将查询-键注意力轨迹和池化隐藏状态解码为相对6自由度位姿。据我们所知,这是首个直接从中间视频扩散表示而非生成像素解码物体6自由度轨迹的方法。DreamTraj在平移和旋转方面均达到了新的最先进水平,超越了使用多帧或特权输入的预测器,并且运行速度比先生成再提取的管道快4.6倍。






RoboTrustBench
Derived from paper: RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
RigidBench
Derived from paper: RigidBench: Evaluating Rigid-Body Physics in Video Generation Models
ChangChrisLiu/GNN_Disassembly_WorldModel
GNN Constraint-Aware World Model Dataset (v3) Real robot episodes with per-frame constraint graphs, SAM2 segmentation masks + 256-D feature embeddings, full 3D depth bundles, and synchronized robot states across two manipulation domains. Both domains share the v3 on-disk layout (same JSON/NPZ schemas, same delta-encoded frame_states, same fully-connected PyG expansion at load time) and now share a unified 270-D node feature format — the PyG loader reads a fixed 10-D type… See the full description on the dataset page: https://huggingface.co/datasets/ChangChrisLiu/GNN_Disassembly_WorldModel.
GroupToM-Bench
Introduces GroupToM-Bench, a multimodal benchmark evaluating group-level theory of mind in MLLMs, revealing gaps in social world modeling.
Kalso42/WorldModelForMaze
WorldModelForMaze Code, datasets, and trained checkpoints for studying world-model representations in maze navigation, based on a modified NanoGPT. Contents *.py — training, testing, probing, and visualization scripts (see readme.md). model/ — architectures: transformer, transformer-rope, transformer-nextlat, mamba, mamba2, gated-deltanet, gru. data/maze/100/ — tokenized maze datasets for Tasks A/C/E/H/I (RWs paths, 100 nodes). out/ — final (10000-iter)… See the full description on the dataset page: https://huggingface.co/datasets/Kalso42/WorldModelForMaze.
Kalso42/WorldModelForMazeWithX
WorldModelForMazeWithX Maze pathfinding sequences for training/probing sequence models (Transformer, Mamba, GRU, Gated-DeltaNet, ...). Task C1: relative-turn navigation on a fixed 10×10 directed grid. Includes a special x terminator marking wall-hit (illegal) paths, used to study a model's ability to recognize its own errors. Maze 10×10 grid, 100 nodes (0–99). Directed edges (down/right, both directions added), edge probability 0.6. Graph:… See the full description on the dataset page: https://huggingface.co/datasets/Kalso42/WorldModelForMazeWithX.
kfallah/world-model-pi-sft-with-cot
world-model-pi-sft-with-cot SFT-ready data for training a world model of the TeichAI pi coding-agent harness. Each world-model assistant target carries a teacher-distilled <think>{rationale}</think> block explaining why the next environment turn follows from prior context, followed by the original [developer] / [tool:*] environment block. Stage 2 of a 2-stage pipeline. Stage 1 dataset (with <COT_PLACEHOLDER> slots) is at kfallah/world-model-pi-sft-formatted. Schema… See the full description on the dataset page: https://huggingface.co/datasets/kfallah/world-model-pi-sft-with-cot.
mental-world-model/menti-bench
Menti-Bench Menti-Bench is a manually constructed, quality-controlled benchmark of situated decision scenarios for evaluating Mental World Modeling (MWM): whether a model can predict what a target agent will actually do next, in scenes where the correct prediction depends on tracking each agent's beliefs, knowledge access, goals, emotions, and social constraints rather than the physical scene alone. Each instance presents a short story (text, an image sequence, or a sounding… See the full description on the dataset page: https://huggingface.co/datasets/mental-world-model/menti-bench.
NTU-yiwen/code-world-model-inference-examples-40
Inference examples This directory contains 40 numbered, independent inference examples. Every example uses only its public number; source case names and internal paths are intentionally omitted. Each numbered directory contains: first_frame.png: exact 1536x864 generated RGB first frame used by inference. prompts/*.txt: the exact rolling long-inference prompts used for the result. condition/*.npz: ordered lossless condition chunks. metadata.json: frame count, FPS, prompt windows… See the full description on the dataset page: https://huggingface.co/datasets/NTU-yiwen/code-world-model-inference-examples-40.
NTU-yiwen/code-world-model-project-page-videos
Code World Model Project Page Videos Public research-demo video assets used by the Code World Model project page. The gallery/ directory contains aligned RGB and proxy videos for interactive comparison.
nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios
Dataset Description: PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios is a large-scale synthetic video dataset of autonomous-driving scenes generated with NVIDIA's internal Omniverse simulation platform. Each clip is a temporally consistent multi-camera surround capture of one ego vehicle and surrounding traffic participants, paired with per-camera VLM captions. The dataset is designed to fill gaps in real-world driving data along two axes: (1) targeted long-tail… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios.
nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes
Dataset Description: The SDG-SynHuman is a large-scale synthetic video dataset of digital humans rendered in diverse indoor and outdoor 3D environments. The dataset contains 236,937 clips, totaling approximately 5,841 hours of video, and is designed to support training and post-training of NVIDIA Cosmos world foundation models and related physical AI research. Each sample is a temporally coherent 60-120 second video clip rendered at 1080p and 30 fps. Clips contain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes.
nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes
PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card Dataset Description PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.
nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card Dataset Description PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding. Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.
nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes
PhysicalAI SDG-Warehouse PhysicalAI SDG-Warehouse is a synthetic, fully-annotated video dataset of staged industrial-safety events captured in a simulated warehouse environment. It contains approximately 123k video clips, totaling roughly 412 hours of footage at 1920x1080 resolution and 30 frames per second, organized across four scenarios: a forklift near-miss with a human worker, a warehouse fire with worker evacuation, a forklift collision with a storage shelf, and a routine… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes.
open-gigaai/CVPR-2026-WorldModel-Track-Dataset
GigaBrain Challenge 2026 (CVPR 2026 Workshop Competition) Registration To access the dataset you must register your team. Required information: Team name Team leader Team members Organization Leader email Click Request Access to participate. Resources After approval you will be able to download: Training dataset Test dataset Baseline model Evaluation scripts
PatronusAI/world_model_corpus
Dataset Card for World Model Corpus The world model corpus contains a set of generated trajectories that are shaped for text-based world modeling task as used by the paper: "Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL". The dataset contains trajectories from nine distinct environments: Tau2Bench, SWE-Smith, DeepresearchQA, Openresearcher, Gorilla/BFCLv4, Webshop, Toolathlon, Pandora and Coderforge. Loading from… See the full description on the dataset page: https://huggingface.co/datasets/PatronusAI/world_model_corpus.
shubhxho/ego-world-model-v1
language: en license: cc-by-4.0 size_categories: 10K<n<100K source_datasets: MicroAGI-Labs/MicroAGI00 task_categories: video-classification image-segmentation depth-estimation object-detection pretty_name: Ego World Model Dataset v1 tags: ego-centric world-model robotics sam3 depth-estimation optical-flow action-conditioned first-person-video--- Ego World Model Dataset v1 Egocentric RGB + metric depth + SAM 3 instance segmentation + action labels, built for training… See the full description on the dataset page: https://huggingface.co/datasets/shubhxho/ego-world-model-v1.
syCen/action-worldmodel-bench
Hugging Face dataset: syCen/action-worldmodel-bench
twelvedata/financial-world-model
Twelve Data World Model Dataset A multi-modal financial time-series dataset built from Twelve Data market data. Each timeframe is published in three parallel views: bars_* — OHLCV bars enriched with causal technical indicators and macro context, in Parquet. text_* — instruction-tuning prompts/labels derived from the bars, in JSONL. trajectories_* — fixed-length rolling windows of state vectors plus next-state pairs, suitable for world-model / sequence-model training, in… See the full description on the dataset page: https://huggingface.co/datasets/twelvedata/financial-world-model.
VideoWorldmodel/ReasoningStructureTestset
Reasoning-Structured Videos A Stratified Diagnostic Suite for Compositional Consistency in Action-Conditioned Video World Models. Reasoning-Structured Videos is a UE5-rendered video benchmark whose trajectories are organised as rooted graphs with path-level algebraic relations. Unlike flat corpora that release independent action–observation rollouts, every released trajectory here is annotated as an exact instance of one of three identities a faithful transition operator must… See the full description on the dataset page: https://huggingface.co/datasets/VideoWorldmodel/ReasoningStructureTestset.
VideoWorldmodel/ReasoningVideoSamples
Hugging Face dataset: VideoWorldmodel/ReasoningVideoSamples

xwk123/Mobile-GUI-Worldmodel-SFT
Mobile-GUI-Worldmodel-SFT This repository contains mobile GUI agent data and auxiliary files for training and evaluating GUI world models. The data is organized around GUI trajectories: each step has a screenshot and page-state annotations such as HTML, plain text, and structured text. Repository Layout . ├── GUI-agent-main/ # Data annotation scripts and examples ├── eval/ # Evaluation assets │ └── AndroidControl_images.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/xwk123/Mobile-GUI-Worldmodel-SFT.
ywang077/Trajectory_world_model_dataset
WestWorld Pretraining Dataset This repository contains the pretraining dataset for WestWorld, a knowledge-encoded scalable trajectory world model for diverse robotic systems. Paper | Project Page | GitHub Description WestWorld is designed to address the scalability challenges in trajectory world models for diverse robotic systems. The dataset includes trajectories from 89 complex environments spanning diverse morphologies across both simulation and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ywang077/Trajectory_world_model_dataset.
ZaidGhazal/world-models-eval
DreamGrasp: Processed LIBERO Manipulation Demonstrations Does a robot policy's evaluation still mean something if it never touched a real simulator, only a world model's imagination of one? This dataset is the shared training data behind that question, a single, ready-to-train release built from LIBERO's manipulation demonstrations (libero_spatial, libero_object, libero_goal). It provides: Fixed, versioned train / validation / test / held-out splits, so every result trained on… See the full description on the dataset page: https://huggingface.co/datasets/ZaidGhazal/world-models-eval.
1x-technologies/world_model_raw_data
Raw Dataset for the 1X World Model Sammpling Challenge. Download with: huggingface-cli download 1x-technologies/worldmodel_raw_data --repo-type dataset --local-dir data Train/Val v2.0 The training dataset is shareded into 100 independent shards. The definitions are as follows: video_{shard}.mp4: Raw video with a resolution of 512x512. segment_idx_{shard}.bin - Maps each frame i to its corresponding segment index. You may want to use this to separate non-contiguous frames from… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_raw_data.
1x-technologies/world_model_tokenized_data
1X World Model Compression Challenge Dataset This repository hosts the dataset for the 1X World Model Compression Challenge. huggingface-cli download 1x-technologies/worldmodel --repo-type dataset --local-dir data Updates Since v1.1 Train/Val v2.0 (~100 hours), replacing v1.1 Test v2.0 dataset for the Compression Challenge Faces blurred for privacy New raw video dataset (CC-BY-NC-SA 4.0) at worldmodel_raw_data Example scripts now split into: cosmos_video_decoder.py —… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_tokenized_data.
Clementppr/lerobot_pick_and_place_dataset_world_model
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 30, "total_frames": 13572, "total_tasks": 1, "total_videos": 30, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:30"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Clementppr/lerobot_pick_and_place_dataset_world_model.


thuml/bytesized32-world-model-cot
See https://github.com/thuml/RLVR-World for examples for using this dataset. Citation @article{wu2025rlvr, title={RLVR-World: Training World Models with Reinforcement Learning}, author={Jialong Wu and Shaofeng Yin and Ningya Feng and Mingsheng Long}, journal={arXiv preprint arXiv:2505.13934}, year={2025}, }
