Awesome World Model Hub 论文 · 数据集 · 项目

数据集大全

数据集

按表示方法和任务分组的数据集。

Preview for IRASim
2024

IRASim

IRASim是一种用于机器人操作的细粒度世界模型,基于动作生成视频,改进了动作-帧对齐并支持规划。

Preview for 迈向交互式视频世界建模
2026

迈向交互式视频世界建模

交互式视频世界建模的综述,涵盖前沿、挑战、基准和未来趋势。

world modelinginteractive video generationaction-conditioned generationsurveybenchmarksCV
Preview for DrivingGen
2026

DrivingGen

DrivingGen是首个生成式驾驶世界模型综合基准,用新指标和多样化数据评估视觉、轨迹、时间一致性和可控性。

Preview for ReWorld
2026

ReWorld

ReWorld利用强化学习通过多维奖励建模将基于视频的具身世界模型与物理真实性、任务逻辑和视觉质量对齐。

Preview for Solaris
2026

Solaris

Solaris提出多人视频世界模型,在Minecraft中实现多智能体一致多视角模拟。

Preview for OSCAR
2026

OSCAR

OSCAR是一个动作条件视频世界模型,通过统一骨架表示和大规模数据,跨机器人形态泛化用于策略评估。

roboticsworld modelvideo predictionembodimentaction conditioninggeneralization
Preview for WBench
2026

WBench

WBench是一个多轮基准,用于评估交互式视频世界模型,涵盖五个维度,包含289个测试用例和22个自动指标。

interactive world modelsbenchmarkmulti-turn evaluationvideo qualityphysics complianceinteraction adherence
Preview for Omni-WorldBench
2026

Omni-WorldBench

Omni-WorldBench用提示套件和智能体指标评估4D世界模型交互响应,揭示18个模型局限。

Preview for WHALE
2024

WHALE

WHALE通过behavior-conditioning和retracing-rollout提升世界模型的泛化性与不确定性估计。

Preview for Physics-IQ 验证
2026

Physics-IQ 验证

审计并改进Physics-IQ基准,用于评估视频生成模型的物理理解,优化提示和评分。

video generationworld modelingphysical understandingbenchmark evaluationvideo generative modelsCV
Preview for ARB4WM
2026

ARB4WM

一个用于评估连续控制中世界模型对抗鲁棒性的统一基准,测试对策略、价值和潜在动态的攻击。

adversarial robustnessworld modelscontinuous controlbenchmarkvisual perturbationsAI
Preview for ARB4WM
2026

ARB4WM

一个用于评估连续控制中世界模型对抗鲁棒性的统一基准,测试对策略、价值和潜在动态的攻击。

Preview for Qwen-RobotWorld技术报告
2026

Qwen-RobotWorld技术报告

Qwen-RobotWorld是语言条件视频世界模型,统一多个具身领域,在多项基准测试中排名第一。

embodied AIworld modelvideo generationlanguage-conditionedroboticsautonomous driving
Preview for stable-worldmodel
2026

stable-worldmodel

一个用于可复现世界模型研究的开源平台,提供标准化数据、基线和评估基准。

world modelsreproducibilityevaluationopen-sourcedata pipelinegeneralization
Preview for SimWorld
2025

SimWorld

提出SimWorld基准,融合模拟引擎与世界模型,可控生成场景以提升自动驾驶感知。

Preview for WorldOlympiad
2026

WorldOlympiad

WorldOlympiad基准测试从物理、几何和交互三个维度评估视频世界模型,覆盖游戏、机器人和真实场景。

world modelsvideo generationbenchmarkphysical faithfulnessgeometric consistencyinteraction fidelity
Preview for 灵巧世界模型
2025

灵巧世界模型

一个视频扩散框架,模拟灵巧人类动作在静态3D场景中引发动态变化,用于交互式数字孪生。

Preview for WorldReasonBench
2026

WorldReasonBench

提出WorldReasonBench基准,测试视频生成器作为世界状态预测器,揭示视觉合理性与推理差距。

video generationbenchmarkworld-state predictionreasoningconsistencyevaluation
Preview for 智能体世界建模
2025

智能体世界建模

提出'levels x laws'世界模型分类法,定义三个能力层级与四个法则领域,综合400余篇文献。

Preview for EchoWorld
2025

EchoWorld

EchoWorld利用运动感知世界模型进行超声心动图探头引导,通过预训练掩码预测和运动感知注意力减少误差。

Preview for Holo-World
2026

Holo-World

Holo-World从单张图像统一控制相机、物体和天气,通过数据集和新颖适配器生成视频世界模型。

video world modelcontrollable video generationcamera controlobject controlweather transferdataset
Preview for 假设世界
2026

假设世界

提出一个因果基准,用于评估视频生成世界模型在物理干预任务上的表现,结果显示所有模型均显著失败。

video generationworld modelscausal reasoningbenchmarkphysical plausibilityembodied scenarios
Preview for 当前世界模型缺乏持久状态核心
2026

当前世界模型缺乏持久状态核心

本文提出WRBench基准,测试世界模型在未被观测时是否维持持久状态,发现当前模型在遮挡期间无法推进事件演化。

world modelsbenchmarkpersistent statecamera motionobservabilityCV
Preview for VerseCrafter
2026

VerseCrafter

VerseCrafter利用4D几何控制(点云和3D高斯轨迹)生成逼真、视角一致的视频,精确控制相机和多物体运动。

Preview for PointWorld
2026

PointWorld

PointWorld是一个大规模预训练的3D世界模型,通过预测3D点流实现跨具身实时MPC操作。

Preview for 机器人操作的世界模型
2026

机器人操作的世界模型

综述机器人操作的世界模型,将其定义为动作条件预测系统,并按表示、预测-动作耦合和使用流程组织方法。

world modelsrobotic manipulationlatent dynamicsaction-conditioned video generationphysics-informed simulationRO
Preview for WorldFly
2026

WorldFly

WorldFly将世界模型与VLA结合用于无人机导航,通过流匹配预测未来视频和动作,在密集城市环境中优于基线方法。

UAV navigationVision-Language-ActionWorld ModelUrban CanyonNavigationAI
Preview for ActWorld
2026

ActWorld

ActWorld通过动作感知记忆和新数据集,将导航世界模型扩展至支持物体交互。

interactive world modelsaction-aware memoryobject interactionnavigationchunk-autoregressiveCV
Preview for Embody4D
2026

Embody4D

Embody4D通过生成式视频到视频世界模型,将单目机器人视频转换为新视角视频,用于具身4D世界建模。

embodied AIworld modelnovel view synthesisvideo generation4D modelingdata engine
Preview for KAN-Dreamer
2025

KAN-Dreamer

将KAN集成到DreamerV3世界模型,在控制任务上实现样本效率和训练速度持平。

Preview for GigaWorld-Policy
2026

GigaWorld-Policy

GigaWorld-Policy是一种以动作为中心的世界-动作模型,高效预测未来动作并可选生成视频,用于机器人策略学习。

Preview for CityBench
2024

CityBench

CityBench系统基准通过交互模拟器评估LLMs和VLMs作为世界模型在13个城市中的城市任务。

Preview for OccDirector
2026

OccDirector

OccDirector 从自然语言生成4D占用动态,实现最先进的指令跟随。

4D occupancyautonomous drivinglanguage-guided generationmulti-agent interactionworld modelvideo generation
Preview for τ₀-WM
2026

τ₀-WM

一种统一视频-动作世界模型,集成策略学习、视频预测与动作评估,用于机器人操作。

region:usworld-modelvideo-generationbenchmarkreproducibility
Preview for WorldSimBench
2024

WorldSimBench

WorldSimBench提出双重评估框架,评估视频生成模型作为世界模拟器,涵盖具身场景。

Preview for Prisma-World
2026

Prisma-World

Prisma-World是一个可通过相机控制的多智能体视频世界模型,通过几何感知去噪和新合成数据集确保跨视角一致性。

video world modelsmulti-agentcamera controlcross-view consistencygeometry-aware denoisingCV
Preview for DrivingDojo数据集
2024

DrivingDojo数据集

DrivingDojo数据集支持具有多样化驾驶操作的交互式世界模型,并提供了动作控制的未来预测基准。

Preview for FieldSeer I
2025

FieldSeer I

几何感知世界模型从部分观测预测电磁场动力学,实现光子设计的交互式数字孪生。

Preview for 3D-VLA
2024

3D-VLA

3D-VLA通过LLM和扩散模型构建生成世界模型,整合三维感知、推理与动作,提升具身规划能力。

Preview for $\mu_0$
2026

$\mu_0$

μ0从视频预测3D交互轨迹,无需动作标签实现可扩展机器人学习。

size_categories:1K<n<10Kmodality:imagemodality:videolibrary:datasetslibrary:mlcroissantregion:us
Preview for WorldMark
2026

WorldMark

WorldMark是首个交互式视频世界模型统一基准,通过标准化场景、动作和评估指标实现公平比较。

benchmarkworld modelsinteractive video generationevaluationaction mappingcomputer vision
Preview for 从文字到世界
2025

从文字到世界

LLM可作为隐式文本世界模型用于智能体强化学习,但收益取决于行为覆盖率和环境复杂性。

Preview for 鲁棒梦想家
2025

鲁棒梦想家

鲁棒梦想家通过潜在高斯记忆与偏差学习解决动作控制视频生成中的漂移问题。

Preview for 修剪视觉世界建模评估的长尾
2026

修剪视觉世界建模评估的长尾

引入Tailor-Bench评估视觉世界模型在长尾物理交互上的表现,揭示其在常见场景之外的泛化能力有限。

computer visionworld modelsbenchmarkphysical reasoninglong-tail distributionevaluation
Preview for PhysicsMind
2026

PhysicsMind

PhysicsMind是一个统一基准,包含真实与模拟环境,用于评估视觉语言模型和世界模型的物理推理与预测。

Preview for PhyGround
2026

PhyGround

PhyGround通过250个提示、13条物理定律和大规模人类研究,对生成式世界模型中的物理推理进行基准测试。

task_categories:text-to-videotask_categories:image-to-videosize_categories:n<1Kformat:jsonmodality:textmodality:video
Preview for DeTrack
2026

DeTrack

用于交互式3D环境中无人机具身跟踪的基准和高度感知双世界模型

drone-embodied trackingaerial object trackingbenchmarkactive perceptionclosed-loop controldual world model
Preview for SmallWorlds
2025

SmallWorlds

提出SmallWorld基准测试,用于在隔离、受控环境中系统评估世界模型的动态理解能力。

Preview for SlowFast-VGen
2024

SlowFast-VGen

SlowFast-VGen提出双速学习,结合慢世界动态和快情景记忆,用于动作驱动长视频生成。

Preview for ReactSim-Bench
2026

ReactSim-Bench

提出ReactSim-Bench,用解耦控制和多种指标评估自动驾驶行为世界模型反应能力。

autonomous drivingbehavior simulationreactive capabilitybenchmarkingworld modelRO
Preview for 4DWorldBench
2025

4DWorldBench

4DWorldBench是一个统一评估3D/4D世界生成模型的框架,衡量感知质量、对齐、物理真实感和一致性。

Preview for Dreamland
2025

Dreamland

Dreamland结合物理模拟器和生成模型,实现可控且逼真的世界创建,提升图像质量和可控性。

Preview for PRISM
2026

PRISM

PRISM从世界模型的冻结编码器中提取动作先验,引导基于模型的规划中的采样,以最小架构开销提升成功率。

world modelsmodel-based planningaction priorcontinuous controlreinforcement learningimagination sampling
Preview for OccProphet
2024

OccProphet

OccProphet利用观察者-预测者-优化者框架预测3D占用,计算量减少58-78%并提升精度。

Preview for NarrativeWorldBench
2026

NarrativeWorldBench

提出长程音频剧基准NarrativeWorldBench和潜世界模型N-VSSM,后者在200集内保持高一致性,超越前沿LLM。

audio dramalong-horizonbenchmarknarrative understandingstate-space modelLLM evaluation
Preview for EgoCS-400K
2026

EgoCS-400K

提出EgoCS-400K,大规模第一人称游戏数据集,含对齐视频-动作-语言轨迹,用于训练交互世界模型。

egocentricdatasetworld modelsCounter-StrikegameplayCV
Preview for EA-WM
2026

EA-WM

EA-WM用结构化运动-视觉动作场和事件感知融合提升世界模型视频生成,在WorldArena达SOTA。

Preview for YoCausal
2026

YoCausal

YoCausal利用反转真实视频评估视频扩散模型的因果理解,揭示时间感知与真实因果的差距。

video diffusion modelsworld modelscausalitybenchmarkcounterfactual reasoningvideo generation
Preview for YoCausal
2024

YoCausal

YoCausal利用反转真实视频评估视频扩散模型的因果理解,揭示时间感知与真实因果的差距。

Preview for Text2World
2025

Text2World

提出Text2World基准,基于PDDL评估LLM生成符号世界模型,发现强化学习训练的推理模型能力仍有限。

Preview for LEIA
2026

LEIA

LEIA是一个用于交互式模拟架构材料的世界模型,能够实时预测变形和应力场。

world modelarchitected materialsmachine learninginteractive simulationfinite element methodmicrostructure
Preview for ReflectiChain
2026

ReflectiChain

ReflectiChain通过生成式供应链世界模型和双环学习弥合LLM与RL差距,在半导体基准上提升韧性。

supply chain resilienceepistemic groundingworld modellarge language modelsreinforcement learninguncertainty quantification
Preview for 世界中的世界
2025

世界中的世界

介绍World-in-World平台,用于在闭环具身环境中基准测试世界模型,揭示可控性和后训练扩展比视觉质量更重要。

Preview for 教视频生成器记住
2026

教视频生成器记住

ReMind通过面向记忆的数据和课程训练激发视频生成器的动态记忆,在STEVO-Bench上达到最优。

video generationworld modelsdynamic memorydiffusion transformersstate evolutionout-of-sight reasoning
Preview for MBench
2026

MBench

MBench是一个通过实体、环境和因果一致性评估视频世界模型记忆能力的基准,使用真实拍摄视频。

video world modelsmemory capabilitybenchmarkevaluationconsistencyCV
Preview for MBench
2026

MBench

MBench是一个通过实体、环境和因果一致性评估视频世界模型记忆能力的基准,使用真实拍摄视频。

Preview for WorldArena 2.0
2026

WorldArena 2.0

WorldArena 2.0 在模态、功能和平台三个维度扩展了具身世界模型基准测试,并包含真实世界评估。

embodied world modelsbenchmarkmultimodalvisuotactileroboticscomputer vision
Preview for JEDI
2025

JEDI

JEDI结合JEPA预测学习与扩散去噪,实现端到端潜在扩散世界模型,在Atari100k上表现优异。

Preview for NORA-1.5
2025

NORA-1.5

NORA-1.5通过流匹配动作专家和基于世界模型的偏好奖励提升VLA模型可靠性,在仿真和真实任务中表现更优。

Preview for FlowMPC
2024

FlowMPC

FlowMPC将流匹配策略与学习的世界模型结合用于测试时规划,提升了操作任务的成功率。

Preview for GigaWorld-1
2026

GigaWorld-1

系统研究世界模型用于机器人策略评估,分析7个视频世界模型和324k次部署。

Preview for SimGen
2024

SimGen

该论文围绕“SimGen: Simulator-conditioned Driving Scene Generation”研究世界模型相关问题。

Preview for SIMMER
2026

SIMMER

SIMMER用符号世界模型评估LLM规划中的潜在故障,发现高故障率,反事实模拟可减少故障。

LLM planninglatent failuresbenchmarkworld modelautonomous agentsCL
Preview for World-R1
2026

World-R1

World-R1通过强化学习在文本到视频生成中施加3D约束,无需架构修改,提升几何一致性。

text-to-video generation3D constraintsreinforcement learningFlow-GRPOgeometric consistencyworld simulation
Preview for IPR-1
2025

IPR-1

IPR-1结合世界模型与VLM策略及PhysCode,在1000+游戏中提升物理推理,超越GPT-5。

Preview for UniOcc
2025

UniOcc

UniOcc是自动驾驶中占用预测与预报的统一基准,整合真实与模拟数据及新颖指标。

Preview for ST-Gen4D
2026

ST-Gen4D

ST-Gen4D将4D时空认知嵌入世界模型,利用图和扩散实现一致的4D生成。

4D generationspatiotemporal cognitionworld modelgenerative modelsdynamic topologyCV
Preview for Drive&Gen
2025

Drive&Gen

提出DriveGen,利用统计度量和合成数据协同评估端到端驾驶与视频生成模型。

Preview for MagicTime
2024

MagicTime

MagicTime通过解耦训练和专用数据集生成编码真实世界物理的延时视频,充当变形模拟器。

Preview for 生成模型理解空间
2026

生成模型理解空间

VEGA-3D将预训练视频扩散模型重新用作潜在世界模拟器,提取隐式3D先验以增强多模态大语言模型的空间推理,无需显式3D监督。

Preview for ShareVerse
2026

ShareVerse

ShareVerse通过空间拼接和跨智能体注意力在CARLA数据上实现多智能体一致性视频生成,用于共享世界建模。

Preview for Orca
2026

Orca

Orca是通用世界基础模型,通过Next-State-Prediction从多模态数据学习统一世界潜在空间,实现文本、图像和动作生成。

Preview for PragWorld
2025

PragWorld

评估LLMs在最小语言变化下对话中局部世界模型的鲁棒性,提出可解释性与微调方法。

Preview for 视频中的思考
2026

视频中的思考

提出CGDJ框架评估视频生成器是否真正推理现实世界,揭示感知-预测差距。

Preview for 动态稀疏性
2025

动态稀疏性

本文挑战了机器人强化学习中世界模型的常见稀疏性假设,发现全局稀疏性罕见,但局部状态依赖的稀疏性常见。

Preview for Deform360
2026

Deform360

大规模多视图触觉数据集,评估可变形物体世界模型,比较2D视频与3D粒子方法。

Preview for PanoWorld
2026

PanoWorld

PanoWorld利用自回归节点生成、3D外壳和高斯溅射缓存,从平面图生成一致的全屋VR全景图。

panoramic video generationworld modelgeometry consistencydepth consistency360-degree videoCV
Preview for 视听世界模型
2025

视听世界模型

提出视听世界模型(AVWM),整合双耳音频与视觉动态,并构建基准和扩散变换器模型用于多模态预测。

Preview for RynnWorld-4D
2026

RynnWorld-4D

RynnWorld-4D从单张RGB-D图像和语言指令生成未来的RGB、深度和光流,用于机器人操作。

Preview for Apple-π
2026

Apple-π

提出Apple-PI基准,通过经典力学任务评估视频模型基于物理定律的推理。

Preview for UniUGP
2025

UniUGP

UniUGP通过混合专家与四阶段训练统一理解、生成与规划,在长尾场景中达到SOTA。

Preview for WorldDiT
2026

WorldDiT

WorldDiT是一种统一的扩散Transformer,用于动作和视觉世界建模,无需大型VLM即可在仿真中取得优异结果。

Preview for DSWorld
2026

DSWorld

DSWorld是数据科学世界模型,预测执行结果,加速RL训练和搜索推理约3-14倍。

Preview for EvolvingWorld
2026

EvolvingWorld

EvolvingWorld提出开放模式框架,实现交互文学中角色与世界模型协同演化,并包含数据集和评估协议。

Preview for VideoCoCo
2026

VideoCoCo

VideoCoCo使用可执行Blender代码作为思维链,通过双引擎框架生成物理一致性视频。

Preview for MemLearner
2026

MemLearner

MemLearner学习自适应查询上下文记忆以用于视频世界模型,在遮挡和动态场景下提升场景一致性。

Preview for PhysMani
2026

PhysMani

PhysMani结合基于物理的三维高斯世界模型与策略,实现动态物体操作,在仿真和真实任务中取得优越成功率。

Preview for GameFactorly
2025

GameFactorly

GameFactory通过多阶段训练与领域适配器,生成开放域动作可控的游戏视频。

Preview for Sekai
2025

Sekai

大规模第一人称视频数据集(5000+小时,100+国家),丰富标注,用于训练世界探索视频生成模型。

Preview for CausalARC
2025

CausalARC

提出CausalARC,一个基于因果世界模型的抽象推理测试平台,并在语言模型上评估。

Preview for RetailSMV
2026

RetailSMV

使用新同步多视角数据集,比较零售场景中基础视频世界模型的自我与外部视角适应。

Preview for MemoBench
2026

MemoBench

MemoBench通过消失-重现范式,使用合成和真实视频片段评估动态环境中的世界建模。

Preview for WorldOdysseyBench
2026

WorldOdysseyBench

提出WorldOdysseyBench,评估交互式世界模型在动作、视觉、物理和记忆维度的长时程稳定性。

Preview for 三思而后行
2025

三思而后行

受世界模型启发的自动驾驶视觉定位框架,通过推理未来空间状态消除自然语言指令歧义。

Preview for OpenSTL
2023

OpenSTL

时空预测学习基准,比较循环与无循环模型在多领域的表现。

Preview for VideoPhy
2024

VideoPhy

VideoPhy基准评估视频生成模型是否遵循真实世界活动的物理常识,发现当前模型严重缺乏。

Preview for VideoVerse
2025

VideoVerse

VideoVerse基准评估T2V模型的时间因果性和世界知识,揭示世界模型能力差距。

Preview for PAI-Bench
2025

PAI-Bench

提出了PAI-Bench,一个通过2,808个真实世界视频案例评估物理人工智能感知与预测的基准。

Preview for EarthNet2021
2021

EarthNet2021

EarthNet2021数据集和挑战,用于基于未来天气预测卫星图像,实现高分辨率地球表面预测。

Preview for MetaOthello
2026

MetaOthello

在多个Othello变体上训练的Transformer共享共同的棋盘状态表示,而非隔离世界模型。

Preview for TC-Bench
2025

TC-Bench

Derived from paper: TC-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation

Preview for STDiff
2023

STDiff

提出STDiff,一种结合神经随机微分方程的时空扩散模型,用于连续随机视频预测,达到最先进性能。

Preview for Thinking Ahead
2025

Thinking Ahead

该论文围绕“Thinking Ahead: Foresight Intelligence in MLLMs and World Models”研究世界模型相关问题。

Preview for AGI Maze
2026

AGI Maze

介绍AGI Maze:通过需记忆和隐藏状态的网格迷宫评估LLMs世界建模的基准。

Preview for 可玩视频生成
2021

可玩视频生成

无监督学习可玩视频生成,用户通过离散动作控制视频,采用自监督编码器-解码器与动作瓶颈。

Preview for PWM-ArtGen
2026

PWM-ArtGen

从单张图像生成铰接3D物体的部分世界模型,联合学习视觉动态和运动学参数

Preview for V-ReasonBench
2025

V-ReasonBench

提出V-ReasonBench基准,用于评估生成模型在四个维度上的视频推理能力,使用合成和真实序列。

Preview for Open-AoE
2026

Open-AoE

Open-AoE提供包含2000小时视频及标注的大规模自我中心操作数据集与工具链,助力具身学习。

Preview for WorldRoamBench
2026

WorldRoamBench

尽管交互式世界模型(IWM)取得了快速进展,现有基准仅在轨迹层面评估动作跟随,忽略了记忆与交互物理。我们提出WorldRoamBench,一个面向长时程稳定性的开放世界基准,涵盖四个维度,每个维度均有定制创新:(i)动作:逐帧动作度量,绕过跨模型语义尺度差异,暴露被轨迹隐藏的失败;(ii)视觉:基于片段的漂移度量,捕捉起始-终点比较无法发现的非单调中期序列崩溃;(iii)物理:可控性门控评估,涵盖力学、光学与3D一致性,在忠实动作执行下对合理性进行评分;(iv)记忆:动作解耦协议,通过过渡定位的3D点云重建评估场景记忆,通过跟踪加VLM推理评估主体记忆。该基准包含600余个测试用例,覆盖自然、城市与室内场景,采用第一/第三人称视角,支持WASD连续交互10-60秒。对10余个开源/闭源模型的评估显示,没有任何模型能可靠满足所有维度;即使最佳模型也仅获得中等分数。在WorldRoamBench上的进步是迈向稳定、物理可信、记忆忠实且可部署于实际应用的交互式世界模型的重要步骤。

Preview for UniVR
2026

UniVR

UniVR利用强化学习从纯视觉演示中学习视觉推理、物理动力学和规划,在新基准上提升25%。

Preview for DreamTraj
2026

DreamTraj

在操作过程中准确预测物体轨迹对于闭环感知-动作循环至关重要。当前进展受限于两个方面:现有数据集缺乏细粒度的语言-运动标注,且现有预测器要么依赖视频、深度或CAD模型等特权输入,要么通过昂贵且易出错的感知管道从完全生成的视频中恢复运动。我们通过MOVE数据集弥补了监督差距,该数据集包含5,038条以物体为中心的自我中心轨迹,每条轨迹均配有细粒度的自然语言指令,而非粗略的动词-名词标签。我们进一步提出了DreamTraj,它仅从单张RGB图像和任务指令即可预测6自由度物体轨迹,推理时无需视频、深度或CAD模型:DreamTraj不生成视频,而是从冻结的图像到视频扩散模型在早期去噪步骤的内部表示中读取运动。一个轻量级的流匹配解码器(Reader)将查询-键注意力轨迹和池化隐藏状态解码为相对6自由度位姿。据我们所知,这是首个直接从中间视频扩散表示而非生成像素解码物体6自由度轨迹的方法。DreamTraj在平移和旋转方面均达到了新的最先进水平,超越了使用多帧或特权输入的预测器,并且运行速度比先生成再提取的管道快4.6倍。

Preview for OSCBench
2026

OSCBench

提出OSCBench基准,评估T2V中对象状态变化;模型在准确状态变换上表现不佳。

Preview for WildCity
2026

WildCity

WildCity是一个真实世界城市规模的多模态数据集和模拟器,用于渲染、仿真和空间智能研究。

Preview for ACE-Data-0
2026

ACE-Data-0

ACE-Data-0:真實家庭中多模態人機交互的數據引擎,推動具身智能。

Preview for RoboTrustBench
2026

RoboTrustBench

Derived from paper: RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

Preview for ChangChrisLiu/GNN_Disassembly_WorldModel
2026

ChangChrisLiu/GNN_Disassembly_WorldModel

GNN Constraint-Aware World Model Dataset (v3) Real robot episodes with per-frame constraint graphs, SAM2 segmentation masks + 256-D feature embeddings, full 3D depth bundles, and synchronized robot states across two manipulation domains. Both domains share the v3 on-disk layout (same JSON/NPZ schemas, same delta-encoded frame_states, same fully-connected PyG expansion at load time) and now share a unified 270-D node feature format — the PyG loader reads a fixed 10-D type… See the full description on the dataset page: https://huggingface.co/datasets/ChangChrisLiu/GNN_Disassembly_WorldModel.

task_categories:roboticstask_categories:image-segmentationtask_categories:graph-mllanguage:enlicense:cc-by-4.0size_categories:1K<n<10K
Preview for GroupToM-Bench
2026

GroupToM-Bench

Introduces GroupToM-Bench, a multimodal benchmark evaluating group-level theory of mind in MLLMs, revealing gaps in social world modeling.

group theory of mindmultimodal large language modelssocial emergencebenchmarkcollective behaviorCV
Preview for Kalso42/WorldModelForMaze
2026

Kalso42/WorldModelForMaze

WorldModelForMaze Code, datasets, and trained checkpoints for studying world-model representations in maze navigation, based on a modified NanoGPT. Contents *.py — training, testing, probing, and visualization scripts (see readme.md). model/ — architectures: transformer, transformer-rope, transformer-nextlat, mamba, mamba2, gated-deltanet, gru. data/maze/100/ — tokenized maze datasets for Tasks A/C/E/H/I (RWs paths, 100 nodes). out/ — final (10000-iter)… See the full description on the dataset page: https://huggingface.co/datasets/Kalso42/WorldModelForMaze.

task_categories:otherlicense:mitregion:usmazeworld-modelsequence-modeling
Preview for Kalso42/WorldModelForMazeWithX
2026

Kalso42/WorldModelForMazeWithX

WorldModelForMazeWithX Maze pathfinding sequences for training/probing sequence models (Transformer, Mamba, GRU, Gated-DeltaNet, ...). Task C1: relative-turn navigation on a fixed 10×10 directed grid. Includes a special x terminator marking wall-hit (illegal) paths, used to study a model's ability to recognize its own errors. Maze 10×10 grid, 100 nodes (0–99). Directed edges (down/right, both directions added), edge probability 0.6. Graph:… See the full description on the dataset page: https://huggingface.co/datasets/Kalso42/WorldModelForMazeWithX.

task_categories:text-generationlanguage:enlicense:mitsize_categories:1M<n<10Mregion:usmaze
Preview for kfallah/world-model-pi-sft-with-cot
2026

kfallah/world-model-pi-sft-with-cot

world-model-pi-sft-with-cot SFT-ready data for training a world model of the TeichAI pi coding-agent harness. Each world-model assistant target carries a teacher-distilled <think>{rationale}</think> block explaining why the next environment turn follows from prior context, followed by the original [developer] / [tool:*] environment block. Stage 2 of a 2-stage pipeline. Stage 1 dataset (with <COT_PLACEHOLDER> slots) is at kfallah/world-model-pi-sft-formatted. Schema… See the full description on the dataset page: https://huggingface.co/datasets/kfallah/world-model-pi-sft-with-cot.

task_categories:text-generationlanguage:enlicense:apache-2.0size_categories:1K<n<10Kregion:usworld-model
Preview for nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios
2026

nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios

Dataset Description: PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios is a large-scale synthetic video dataset of autonomous-driving scenes generated with NVIDIA's internal Omniverse simulation platform. Each clip is a temporally consistent multi-camera surround capture of one ego vehicle and surrounding traffic participants, paired with per-camera VLM captions. The dataset is designed to fill gaps in real-world driving data along two axes: (1) targeted long-tail… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios.

language:enlicense:othersize_categories:100K<n<1Mmodality:videoregion:usvideo
Preview for nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes
2026

nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes

Dataset Description: The SDG-SynHuman is a large-scale synthetic video dataset of digital humans rendered in diverse indoor and outdoor 3D environments. The dataset contains 236,937 clips, totaling approximately 5,841 hours of video, and is designed to support training and post-training of NVIDIA Cosmos world foundation models and related physical AI research. Each sample is a temporally coherent 60-120 second video clip rendered at 1080p and 30 fps. Clips contain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes.

language:enlicense:othermodality:videoregion:usvideosynthetic
Preview for nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes
2026

nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes

PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card Dataset Description PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.

license:othersize_categories:100K<n<1Mmodality:videoregion:usroboticssynthetic-data
Preview for nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes
2026

nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes

PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card Dataset Description PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding. Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.

license:othersize_categories:100M<n<1Bformat:webdatasetmodality:imagemodality:textlibrary:datasets
Preview for nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes
2026

nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes

PhysicalAI SDG-Warehouse PhysicalAI SDG-Warehouse is a synthetic, fully-annotated video dataset of staged industrial-safety events captured in a simulated warehouse environment. It contains approximately 123k video clips, totaling roughly 412 hours of footage at 1920x1080 resolution and 30 frames per second, organized across four scenarios: a forklift near-miss with a human worker, a warehouse fire with worker evacuation, a forklift collision with a storage shelf, and a routine… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes.

task_categories:video-classificationtask_categories:video-text-to-texttask_categories:text-to-videolanguage:enlicense:othersize_categories:100K<n<1M
Preview for open-gigaai/CVPR-2026-WorldModel-Track-Dataset
2026

open-gigaai/CVPR-2026-WorldModel-Track-Dataset

GigaBrain Challenge 2026 (CVPR 2026 Workshop Competition) Registration To access the dataset you must register your team. Required information: Team name Team leader Team members Organization Leader email Click Request Access to participate. Resources After approval you will be able to download: Training dataset Test dataset Baseline model Evaluation scripts

region:us
Preview for PatronusAI/world_model_corpus
2026

PatronusAI/world_model_corpus

Dataset Card for World Model Corpus The world model corpus contains a set of generated trajectories that are shaped for text-based world modeling task as used by the paper: "Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL". The dataset contains trajectories from nine distinct environments: Tau2Bench, SWE-Smith, DeepresearchQA, Openresearcher, Gorilla/BFCLv4, Webshop, Toolathlon, Pandora and Coderforge. Loading from… See the full description on the dataset page: https://huggingface.co/datasets/PatronusAI/world_model_corpus.

task_categories:text-generationlicense:mitsize_categories:100K<n<1Mformat:parquetmodality:textlibrary:datasets
Preview for shubhxho/ego-world-model-v1
2026

shubhxho/ego-world-model-v1

language: en license: cc-by-4.0 size_categories: 10K<n<100K source_datasets: MicroAGI-Labs/MicroAGI00 task_categories: video-classification image-segmentation depth-estimation object-detection pretty_name: Ego World Model Dataset v1 tags: ego-centric world-model robotics sam3 depth-estimation optical-flow action-conditioned first-person-video--- Ego World Model Dataset v1 Egocentric RGB + metric depth + SAM 3 instance segmentation + action labels, built for training… See the full description on the dataset page: https://huggingface.co/datasets/shubhxho/ego-world-model-v1.

region:us
Preview for twelvedata/financial-world-model
2026

twelvedata/financial-world-model

Twelve Data World Model Dataset A multi-modal financial time-series dataset built from Twelve Data market data. Each timeframe is published in three parallel views: bars_* — OHLCV bars enriched with causal technical indicators and macro context, in Parquet. text_* — instruction-tuning prompts/labels derived from the bars, in JSONL. trajectories_* — fixed-length rolling windows of state vectors plus next-state pairs, suitable for world-model / sequence-model training, in… See the full description on the dataset page: https://huggingface.co/datasets/twelvedata/financial-world-model.

task_categories:time-series-forecastingtask_categories:text-generationtask_categories:reinforcement-learninglanguage:enlicense:mitsize_categories:10M<n<100M
Preview for VideoWorldmodel/ReasoningStructureTestset
2026

VideoWorldmodel/ReasoningStructureTestset

Reasoning-Structured Videos A Stratified Diagnostic Suite for Compositional Consistency in Action-Conditioned Video World Models. Reasoning-Structured Videos is a UE5-rendered video benchmark whose trajectories are organised as rooted graphs with path-level algebraic relations. Unlike flat corpora that release independent action–observation rollouts, every released trajectory here is annotated as an exact instance of one of three identities a faithful transition operator must… See the full description on the dataset page: https://huggingface.co/datasets/VideoWorldmodel/ReasoningStructureTestset.

task_categories:video-classificationtask_categories:otherlanguage:enlicense:cc-by-4.0size_categories:1K<n<10Kregion:us
Preview for Wh0
2026

Wh0

Wh0 uses generative video world models to produce scalable egocentric human-hand manipulation data, improving dexterous VLA model zero-shot success on real-world tasks.

generative world modelsegocentric videodexterous manipulationdata generationhuman-object interactionVLA models
Preview for xwk123/Mobile-GUI-Worldmodel-SFT
2026

xwk123/Mobile-GUI-Worldmodel-SFT

Mobile-GUI-Worldmodel-SFT This repository contains mobile GUI agent data and auxiliary files for training and evaluating GUI world models. The data is organized around GUI trajectories: each step has a screenshot and page-state annotations such as HTML, plain text, and structured text. Repository Layout . ├── GUI-agent-main/ # Data annotation scripts and examples ├── eval/ # Evaluation assets │ └── AndroidControl_images.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/xwk123/Mobile-GUI-Worldmodel-SFT.

task_categories:image-to-texttask_categories:visual-question-answeringtask_categories:roboticslanguage:enlicense:othersize_categories:10K<n<100K
Preview for ywang077/Trajectory_world_model_dataset
2026

ywang077/Trajectory_world_model_dataset

WestWorld Pretraining Dataset This repository contains the pretraining dataset for WestWorld, a knowledge-encoded scalable trajectory world model for diverse robotic systems. Paper | Project Page | GitHub Description WestWorld is designed to address the scalability challenges in trajectory world models for diverse robotic systems. The dataset includes trajectories from 89 complex environments spanning diverse morphologies across both simulation and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ywang077/Trajectory_world_model_dataset.

task_categories:roboticslicense:cc-by-nc-4.0arxiv:2603.14392region:us
Preview for ZaidGhazal/world-models-eval
2026

ZaidGhazal/world-models-eval

DreamGrasp: Processed LIBERO Manipulation Demonstrations Does a robot policy's evaluation still mean something if it never touched a real simulator, only a world model's imagination of one? This dataset is the shared training data behind that question, a single, ready-to-train release built from LIBERO's manipulation demonstrations (libero_spatial, libero_object, libero_goal). It provides: Fixed, versioned train / validation / test / held-out splits, so every result trained on… See the full description on the dataset page: https://huggingface.co/datasets/ZaidGhazal/world-models-eval.

task_categories:roboticslicense:mitsize_categories:100K<n<1Mformat:parquetmodality:tabularmodality:timeseries
Preview for 1x-technologies/world_model_raw_data
2025

1x-technologies/world_model_raw_data

Raw Dataset for the 1X World Model Sammpling Challenge. Download with: huggingface-cli download 1x-technologies/worldmodel_raw_data --repo-type dataset --local-dir data Train/Val v2.0 The training dataset is shareded into 100 independent shards. The definitions are as follows: video_{shard}.mp4: Raw video with a resolution of 512x512. segment_idx_{shard}.bin - Maps each frame i to its corresponding segment index. You may want to use this to separate non-contiguous frames from… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_raw_data.

license:cc-by-nc-sa-4.0size_categories:10M<n<100Mregion:us
Preview for 1x-technologies/world_model_tokenized_data
2025

1x-technologies/world_model_tokenized_data

1X World Model Compression Challenge Dataset This repository hosts the dataset for the 1X World Model Compression Challenge. huggingface-cli download 1x-technologies/worldmodel --repo-type dataset --local-dir data Updates Since v1.1 Train/Val v2.0 (~100 hours), replacing v1.1 Test v2.0 dataset for the Compression Challenge Faces blurred for privacy New raw video dataset (CC-BY-NC-SA 4.0) at worldmodel_raw_data Example scripts now split into: cosmos_video_decoder.py —… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_tokenized_data.

license:apache-2.0size_categories:10M<n<100Mregion:us
Preview for Clementppr/lerobot_pick_and_place_dataset_world_model
2025

Clementppr/lerobot_pick_and_place_dataset_world_model

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 30, "total_frames": 13572, "total_tasks": 1, "total_videos": 30, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:30"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Clementppr/lerobot_pick_and_place_dataset_world_model.

task_categories:roboticslicense:apache-2.0size_categories:10K<n<100Kformat:parquetmodality:tabularmodality:timeseries
Preview for GEM
2025

GEM

GEM is a multimodal world model for controllable future frame prediction with ego-motion, object dynamics, and human pose control, using a large real-world dataset.

Preview for ORV
2025

ORV

ORV introduces a 4D occupancy-centric framework for controllable robot video generation, improving fidelity, temporal consistency, and control alignment.

Preview for thuml/bytesized32-world-model-cot
2025

thuml/bytesized32-world-model-cot

See https://github.com/thuml/RLVR-World for examples for using this dataset. Citation @article{wu2025rlvr, title={RLVR-World: Training World Models with Reinforcement Learning}, author={Jialong Wu and Shaofeng Yin and Ningya Feng and Mingsheng Long}, journal={arXiv preprint arXiv:2505.13934}, year={2025}, }

license:mitsize_categories:100K<n<1Mformat:parquetmodality:textlibrary:datasetslibrary:dask
Preview for EVA
2024

EVA

Proposes EVA, an embodied world model using vision-language and video generation models for multi-step video prediction and OOD handling.