
Toward Memory-Aided World Models
Proposes LoopNav, a Minecraft dataset and benchmark for evaluating spatial consistency in world models using loop-based navigation.
Dataset zoo
Datasets grouped by representation and task.

Proposes LoopNav, a Minecraft dataset and benchmark for evaluating spatial consistency in world models using loop-based navigation.


An open-source benchmark and baseline model for evaluating action fidelity in world models for autonomous driving.

A comprehensive survey on interactive world modeling, covering trends, challenges, benchmarks, and future directions in action-conditioned video/3D generation.

DrivingGen is the first comprehensive benchmark for generative driving world models, with new metrics and diverse data to evaluate visual realism, trajectory plausibility, temporal coherence, and controllability.




A latent Diffusion Transformer world model predicts future AV camera scenes from actions, outperforming regression on perception metrics.

Introduces Hybrid Memory and HyDRA for video world models to maintain subject consistency when hidden, with a new dataset HM-World.

Uses action-conditional video generation models as world models for scalable policy evaluation in robotic manipulation.


Omni-WorldBench evaluates world models' interactive response in 4D settings using a prompt suite and agent-based metrics, revealing limitations across 18 models.

Introduces WoW-World-Eval, a benchmark with 22 metrics to evaluate video foundation models as embodied world models, showing poor performance in planning and physical consistency.

LLMs are tested as text-based world simulators using a new benchmark; GPT-4 proves unreliable, highlighting limitations.


Dream4Drive uses a driving world model to generate synthetic multi-view videos, enhancing downstream perception tasks and corner case detection in autonomous driving.

Audits and improves Physics-IQ benchmark for evaluating physical understanding of video generative models, refining samples and prompts.



A language-conditioned video world model unifying embodied intelligence across robotics, driving, navigation, and human-to-robot transfer.

An open-source platform for standardized, reproducible world modeling research with high-performance data handling, baseline implementations, and systematic evaluation benchmarks.


Proposes WorldModelBench to evaluate video generation models as world models, focusing on physics adherence and instruction-following with human labels.

Compares reconstruction vs semantic latent spaces for action-conditioned video diffusion world models, finding semantic encoders better for policy-relevant robotics.


A benchmark evaluating video-based world models across physical, geometric, and interaction fidelity with real-world scenarios.

A video diffusion framework that models dexterous human actions inducing dynamic changes in static 3D scenes for interactive digital twins.

Introduces WorldReasonBench, a benchmark to evaluate video generators as world-state predictors, revealing a gap between visual quality and reasoning.



Holo-World enables unified camera, object, and weather control from a single image for video generation.

Introduces What-If World, a causal benchmark using prompt pairs to test if video world models correctly respond to physical changes.

This paper introduces WRBench to test if world models maintain persistent internal state when unobserved, finding current models fail to evolve events during occlusion.

Proposes PhyGenesis, a driving video world model that generates physically consistent videos under challenging trajectories using a physical condition generator and physics-enhanced video generator.

VerseCrafter uses 4D geometric control (point clouds and 3D Gaussian trajectories) to generate realistic, view-consistent videos with precise camera and multi-object motion.

Generative World Explorer enables agents to mentally explore 3D worlds and update beliefs for better planning without physical exploration.

PointWorld is a large pre-trained 3D world model that forecasts 3D point flows from RGB-D images and actions, enabling real-time MPC for robotic manipulation across embodiments.

Survey of world models for robotic manipulation, categorizing representations, action connections, and usage in robot learning pipelines.

Introduces world stability metric for diffusion-based world models, evaluates state-of-the-art models, and proposes improvement strategies.


WorldLens benchmark evaluates driving world models across pixel quality, geometry, closed-loop driving, and human perception, revealing no model excels universally.




Integrates KANs into DreamerV3 world model, achieving parity in sample efficiency and training speed on a control task.

GigaWorld-Policy is an action-centered World-Action Model that efficiently predicts future actions and optionally generates videos for robot policy learning.


OccDirector generates 4D occupancy dynamics from natural language for autonomous driving simulation, achieving state-of-the-art instruction-following.

Matrix-Game 2.0 Self-Consistency Reproduction This README explains how to reproduce the Matrix-Game 2.0 self-consistency (SC) row in Reasoning-Structured Videos: A Stratified Diagnostic Suite for Compositional Consistency in World Models. All experimental data, generated videos, and evaluation files required for this reproduction are available on the current dataset page under Files and versions: https://huggingface.co/datasets/VideoWorldmodel/Evaluation No external dataset or… See the full description on the dataset page: https://huggingface.co/datasets/VideoWorldmodel/Evaluation.

WorldSimBench proposes a dual evaluation framework for video generation models as world simulators, covering embodied scenarios.

Prisma-World generates consistent multi-agent videos from multiple camera viewpoints using geometry-aware denoising and cross-view attention.

DrivingDojo dataset enables interactive world models with diverse driving maneuvers and action-controlled future prediction benchmark.

Introduces geometry-grounded long-term spatial memory to improve consistency in video world models.

Geometry-aware world model forecasts electromagnetic field dynamics from partial observations, enabling interactive digital twins for photonic design.

First large-scale video prediction model for autonomous driving using web data and latent diffusion, achieving zero-shot generalization and adaptation to planning.

Hallucination in world models stems from low data coverage; detectable via three signals and preventable with coverage-aware sampling and curiosity-driven finetuning.


Hugging Face dataset: qsun2001/world_model

Introduces ORCA, a closed-loop world modeling framework for video avatars to achieve autonomous goal-directed behavior via hierarchical reasoning and belief updating.

LLMs can serve as implicit text-based world models for agentic RL, but benefits depend on behavioral coverage and environment complexity.


MLLMs like GPT-4o struggle to understand dynamic driving scenes across frames, despite excelling at single images, highlighting gaps in world model capabilities.

Introduces Tailor-Bench to evaluate visual world models on long-tail physical interactions, revealing performance gaps in generalization.

PhysicsMind is a unified benchmark with real and simulated environments to evaluate physical reasoning and prediction in VLMs and world models.

PhyGround Project Page | GitHub | Paper PhyGround is a criteria-grounded benchmark for evaluating physical reasoning in video generation. The benchmark contains 250 curated prompts, each augmented with an expected physical outcome, and a taxonomy of 13 physical laws across solid-body mechanics, fluid dynamics, and optics. Sample Usage You can download the benchmark prompts and first-frame images using the Hugging Face CLI: huggingface-cli download --repo-type dataset… See the full description on the dataset page: https://huggingface.co/datasets/NU-World-Model-Embodied-AI/phyground.


Demonstrates novel data poisoning attacks on world models that compromise robot learning pipelines, generating dangerous synthetic trajectories.

Introduces SmallWorld Benchmark for systematically evaluating world models' dynamics understanding in isolated, controlled environments.

UWM couples video and action diffusion in a unified transformer for pretraining on large robotic datasets, enabling policy and dynamics learning.

SlowFast-VGen introduces dual-speed learning combining slow world dynamics and fast episodic memory for action-driven long video generation.

Introduces ReactSim-Bench, a benchmark for evaluating reactive behavior world model simulation in autonomous driving with decoupled control and safety metrics.

4DWorldBench is a unified evaluation framework for 3D/4D world generation models, assessing perceptual quality, alignment, physical realism, and consistency.

A pilot study evaluating zero-shot surgical video generation, revealing a plausibility gap between visual quality and causal understanding.

Introduces reference-free metrics using DROID-SLAM and SEA-RAFT to evaluate physical consistency in world model-based video generation, improving task success rates by 8%.




Introduces a benchmark and latent world model for long-horizon co-creative audio drama, outperforming frontier LLMs in consistency and controllability.

EgoCS-400K is a large-scale egocentric Counter-Strike dataset with temporally aligned video-action-language trajectories for training interactive world models.

WorldArena benchmark evaluates embodied world models on perception and functional utility, revealing a gap between visual quality and task capability.


Video generation models fail to learn true physical laws, showing case-based generalization and poor out-of-distribution extrapolation.

Introduces Text2World, a PDDL-based benchmark for evaluating LLMs on symbolic world model generation, revealing limited capabilities despite RL-trained reasoning models.


ReflectiChain bridges LLM and RL gaps with a generative supply chain world model and double-loop learning, improving resilience under uncertainty.

Introduces World-in-World, a platform for benchmarking world models in closed-loop embodied environments, revealing that controllability and post-training scaling matter more than visual quality.

Distills game code world model generation into lightweight LLMs via SFT and RLVR, improving syntactic correctness and rule adherence.

Reports on a challenge for holistic quality assessment of world-model-generated videos, evaluating physical realism, temporal consistency, and anomaly localization.

Introduces PDI-Bench, a quantitative framework to evaluate geometric coherence in generated videos, revealing failure modes missed by perceptual metrics.

ReMind elicits dynamic memory in video generators via memory-oriented data and curriculum training, achieving state-of-the-art on STEVO-Bench.

A comprehensive survey proposing a unified framework and taxonomy for world models in embodied AI, covering functionality, temporal modeling, spatial representation, and open challenges.



WorldArena 2.0 expands embodied world model benchmarking across modality, functionality, and platform, including real-world robotic evaluations.




Systematic study of world models for robot policy evaluation using WMBench benchmark, analyzing 7 video world models and 324k rollouts.

Investigates how mobile world models guide GUI agents by comparing four modalities and evaluating downstream utility on benchmarks.

WorldBench is a video benchmark for disentangled evaluation of world models' understanding of individual physics concepts, revealing failures in SOTA models.

Introduces Nuplan-Occ dataset and a unified framework for generating semantic occupancy, multi-view videos, and LiDAR point clouds for autonomous driving.


Proposes world model-based video generation and temporal reasoning to improve accident anticipation in autonomous driving, with a new benchmark dataset.

Proposes WorldTest protocol and AutumnBench benchmark to evaluate world models on multiple environment-level queries, showing humans outperform frontier models.

Introduces ORAD-3D, the largest off-road autonomous driving dataset with benchmarks including a world model task.





Learns particle-based dynamics from real-world videos using Gaussian splatting, avoiding synthetic data reliance.

A survey on multimodal large language models for autonomous driving, covering background, tools, datasets, benchmarks, and future challenges.

LLM-based agents use dialogue to align world models for coordination, but it reduces task success despite fewer conflicts.





Compares robustness of World Action Models (WAMs) vs. VLAs on robotic benchmarks under visual/language perturbations, finding WAMs more robust.

GLIMO uses proxy world models (simulators) to generate training data via an LLM agent with self-refinement, improving LLM performance on embodied tasks.

First multiplayer world model using latent diffusion, trained on 10k hours of Rocket League, generates stable 4-player matches in real time.

ShareVerse enables multi-agent consistent video generation for shared world modeling using spatial concatenation and cross-agent attention on CARLA data.



Proposes Causal-Generative Dual-Judge (CGDJ) to evaluate if video generators truly reason about real-world dynamics, revealing a perception-prediction gap.

This paper challenges common sparsity assumptions in world models for robotic RL, finding global sparsity rare but local state-dependent sparsity common.



Proposes Audio-Visual World Models (AVWM) integrating binaural audio and visual dynamics, with a benchmark and a diffusion transformer model for multimodal prediction.

RynnWorld-4D generates future RGB, depth, and optical flow from a single RGB-D image and language instruction for robotic manipulation.





EvolvingWorld introduces an open-schema framework for co-evolving characters and world models in interactive literary simulations, with a dataset and evaluation protocol.


MemLearner learns to query context memory adaptively for video world models, improving scene consistency under occlusions and dynamics.

World models are a powerful paradigm in AI and robotics, enabling agents to reason about the future by predicting visual observations or compact latent states.

Introduces RBench benchmark and RoVid-X dataset for robot-oriented video generation, revealing deficiencies in physical realism.


GameFactory generates open-domain action-controllable game videos using a multi-phase training strategy with domain adapter.

Proposes a control-centric benchmark (VP^2) for action-conditioned video prediction to evaluate models for robotic manipulation planning.




Paper presents 1M-context language and video-language models using RingAttention, with open-source implementation and benchmarks.


Introduces WorldOdysseyBench, a benchmark for long-horizon stability of interactive world models across action, vision, physics, and memory dimensions.

A world model-inspired framework for autonomous vehicle visual grounding that reasons about future spatial states to disambiguate natural-language commands.


Proposes Large Emotional World Model integrating emotion into world models, improving prediction of emotion-driven social behaviors.


VideoVerse benchmark evaluates T2V models on temporal causality and world knowledge, revealing gaps in world model capabilities.


EarthNet2021 dataset and challenge for forecasting satellite images conditioned on future weather, enabling high-resolution Earth surface predictions.

Transformers trained on multiple Othello variants share a common board-state representation rather than isolating world models.
Derived from paper: TC-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation


Introduces Foresight Intelligence and FSU-QA dataset to evaluate VLMs and world models on reasoning about future events.

Introduces AGI Maze, a benchmark for evaluating world-modeling in LLMs via grid-based mazes requiring memory and hidden state representation.

Unsupervised learning of playable video generation where user controls video by selecting discrete actions, using self-supervised encoder-decoder with action bottleneck.

A part world model for articulated 3D object generation from a single image, learning joint visual dynamics and kinematic parameters.

Introduces V-ReasonBench, a benchmark for evaluating video reasoning in generative models across four dimensions using synthetic and real-world sequences.


Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignore memory and interaction physics. We introduce WorldRoamBench, an open-world benchmark for long-horiz...

Proposes city-specific world models and an MPC planner (City-Driver) for urban navigation, achieving state-of-the-art on nuPlan benchmark.




ACE-Data-0 presents a data engine capturing multimodal human-centric interactions in real homes for embodied intelligence.

Derived from paper: RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
GNN Constraint-Aware World Model Dataset (v3) Real robot episodes with per-frame constraint graphs, SAM2 segmentation masks + 256-D feature embeddings, full 3D depth bundles, and synchronized robot states across two manipulation domains. Both domains share the v3 on-disk layout (same JSON/NPZ schemas, same delta-encoded frame_states, same fully-connected PyG expansion at load time) and now share a unified 270-D node feature format — the PyG loader reads a fixed 10-D type… See the full description on the dataset page: https://huggingface.co/datasets/ChangChrisLiu/GNN_Disassembly_WorldModel.
Introduces GroupToM-Bench, a multimodal benchmark evaluating group-level theory of mind in MLLMs, revealing gaps in social world modeling.
WorldModelForMaze Code, datasets, and trained checkpoints for studying world-model representations in maze navigation, based on a modified NanoGPT. Contents *.py — training, testing, probing, and visualization scripts (see readme.md). model/ — architectures: transformer, transformer-rope, transformer-nextlat, mamba, mamba2, gated-deltanet, gru. data/maze/100/ — tokenized maze datasets for Tasks A/C/E/H/I (RWs paths, 100 nodes). out/ — final (10000-iter)… See the full description on the dataset page: https://huggingface.co/datasets/Kalso42/WorldModelForMaze.
WorldModelForMazeWithX Maze pathfinding sequences for training/probing sequence models (Transformer, Mamba, GRU, Gated-DeltaNet, ...). Task C1: relative-turn navigation on a fixed 10×10 directed grid. Includes a special x terminator marking wall-hit (illegal) paths, used to study a model's ability to recognize its own errors. Maze 10×10 grid, 100 nodes (0–99). Directed edges (down/right, both directions added), edge probability 0.6. Graph:… See the full description on the dataset page: https://huggingface.co/datasets/Kalso42/WorldModelForMazeWithX.
world-model-pi-sft-with-cot SFT-ready data for training a world model of the TeichAI pi coding-agent harness. Each world-model assistant target carries a teacher-distilled <think>{rationale}</think> block explaining why the next environment turn follows from prior context, followed by the original [developer] / [tool:*] environment block. Stage 2 of a 2-stage pipeline. Stage 1 dataset (with <COT_PLACEHOLDER> slots) is at kfallah/world-model-pi-sft-formatted. Schema… See the full description on the dataset page: https://huggingface.co/datasets/kfallah/world-model-pi-sft-with-cot.
Dataset Description: PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios is a large-scale synthetic video dataset of autonomous-driving scenes generated with NVIDIA's internal Omniverse simulation platform. Each clip is a temporally consistent multi-camera surround capture of one ego vehicle and surrounding traffic participants, paired with per-camera VLM captions. The dataset is designed to fill gaps in real-world driving data along two axes: (1) targeted long-tail… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios.
Dataset Description: The SDG-SynHuman is a large-scale synthetic video dataset of digital humans rendered in diverse indoor and outdoor 3D environments. The dataset contains 236,937 clips, totaling approximately 5,841 hours of video, and is designed to support training and post-training of NVIDIA Cosmos world foundation models and related physical AI research. Each sample is a temporally coherent 60-120 second video clip rendered at 1080p and 30 fps. Clips contain… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Digital-Human-Scenes.
PhysicalAI WorldModel Synthetic Embodied Robot Scenes Dataset Card Dataset Description PhysicalAI WorldModel Synthetic Embodied Robot Scenes is a large-scale synthetic robotics video corpus generated from USD-based robotic simulation and rendering pipelines built around NVIDIA Isaac Sim, Omniverse, Isaac Lab, and related robot data-generation systems. It is designed to improve physical plausibility, embodiment persistence, task-conditioned robot behavior reasoning… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card Dataset Description PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes is a large-scale synthetic dataset of physically-simulated multi-object interaction scenes, generated using NVIDIA Isaac Sim and the PhysX physics engine. It is designed to train and evaluate AI models on physical reasoning, rigid body dynamics, optical flow, depth estimation, and scene understanding. Each clip… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes.
PhysicalAI SDG-Warehouse PhysicalAI SDG-Warehouse is a synthetic, fully-annotated video dataset of staged industrial-safety events captured in a simulated warehouse environment. It contains approximately 123k video clips, totaling roughly 412 hours of footage at 1920x1080 resolution and 30 frames per second, organized across four scenarios: a forklift near-miss with a human worker, a warehouse fire with worker evacuation, a forklift collision with a storage shelf, and a routine… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Warehouse-Operations-Scenes.
GigaBrain Challenge 2026 (CVPR 2026 Workshop Competition) Registration To access the dataset you must register your team. Required information: Team name Team leader Team members Organization Leader email Click Request Access to participate. Resources After approval you will be able to download: Training dataset Test dataset Baseline model Evaluation scripts
Dataset Card for World Model Corpus The world model corpus contains a set of generated trajectories that are shaped for text-based world modeling task as used by the paper: "Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL". The dataset contains trajectories from nine distinct environments: Tau2Bench, SWE-Smith, DeepresearchQA, Openresearcher, Gorilla/BFCLv4, Webshop, Toolathlon, Pandora and Coderforge. Loading from… See the full description on the dataset page: https://huggingface.co/datasets/PatronusAI/world_model_corpus.
language: en license: cc-by-4.0 size_categories: 10K<n<100K source_datasets: MicroAGI-Labs/MicroAGI00 task_categories: video-classification image-segmentation depth-estimation object-detection pretty_name: Ego World Model Dataset v1 tags: ego-centric world-model robotics sam3 depth-estimation optical-flow action-conditioned first-person-video--- Ego World Model Dataset v1 Egocentric RGB + metric depth + SAM 3 instance segmentation + action labels, built for training… See the full description on the dataset page: https://huggingface.co/datasets/shubhxho/ego-world-model-v1.
Hugging Face dataset: syCen/action-worldmodel-bench
Twelve Data World Model Dataset A multi-modal financial time-series dataset built from Twelve Data market data. Each timeframe is published in three parallel views: bars_* — OHLCV bars enriched with causal technical indicators and macro context, in Parquet. text_* — instruction-tuning prompts/labels derived from the bars, in JSONL. trajectories_* — fixed-length rolling windows of state vectors plus next-state pairs, suitable for world-model / sequence-model training, in… See the full description on the dataset page: https://huggingface.co/datasets/twelvedata/financial-world-model.
Reasoning-Structured Videos A Stratified Diagnostic Suite for Compositional Consistency in Action-Conditioned Video World Models. Reasoning-Structured Videos is a UE5-rendered video benchmark whose trajectories are organised as rooted graphs with path-level algebraic relations. Unlike flat corpora that release independent action–observation rollouts, every released trajectory here is annotated as an exact instance of one of three identities a faithful transition operator must… See the full description on the dataset page: https://huggingface.co/datasets/VideoWorldmodel/ReasoningStructureTestset.
Hugging Face dataset: VideoWorldmodel/ReasoningVideoSamples

Mobile-GUI-Worldmodel-SFT This repository contains mobile GUI agent data and auxiliary files for training and evaluating GUI world models. The data is organized around GUI trajectories: each step has a screenshot and page-state annotations such as HTML, plain text, and structured text. Repository Layout . ├── GUI-agent-main/ # Data annotation scripts and examples ├── eval/ # Evaluation assets │ └── AndroidControl_images.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/xwk123/Mobile-GUI-Worldmodel-SFT.
WestWorld Pretraining Dataset This repository contains the pretraining dataset for WestWorld, a knowledge-encoded scalable trajectory world model for diverse robotic systems. Paper | Project Page | GitHub Description WestWorld is designed to address the scalability challenges in trajectory world models for diverse robotic systems. The dataset includes trajectories from 89 complex environments spanning diverse morphologies across both simulation and real-world… See the full description on the dataset page: https://huggingface.co/datasets/ywang077/Trajectory_world_model_dataset.
DreamGrasp: Processed LIBERO Manipulation Demonstrations Does a robot policy's evaluation still mean something if it never touched a real simulator, only a world model's imagination of one? This dataset is the shared training data behind that question, a single, ready-to-train release built from LIBERO's manipulation demonstrations (libero_spatial, libero_object, libero_goal). It provides: Fixed, versioned train / validation / test / held-out splits, so every result trained on… See the full description on the dataset page: https://huggingface.co/datasets/ZaidGhazal/world-models-eval.
Raw Dataset for the 1X World Model Sammpling Challenge. Download with: huggingface-cli download 1x-technologies/worldmodel_raw_data --repo-type dataset --local-dir data Train/Val v2.0 The training dataset is shareded into 100 independent shards. The definitions are as follows: video_{shard}.mp4: Raw video with a resolution of 512x512. segment_idx_{shard}.bin - Maps each frame i to its corresponding segment index. You may want to use this to separate non-contiguous frames from… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_raw_data.
1X World Model Compression Challenge Dataset This repository hosts the dataset for the 1X World Model Compression Challenge. huggingface-cli download 1x-technologies/worldmodel --repo-type dataset --local-dir data Updates Since v1.1 Train/Val v2.0 (~100 hours), replacing v1.1 Test v2.0 dataset for the Compression Challenge Faces blurred for privacy New raw video dataset (CC-BY-NC-SA 4.0) at worldmodel_raw_data Example scripts now split into: cosmos_video_decoder.py —… See the full description on the dataset page: https://huggingface.co/datasets/1x-technologies/world_model_tokenized_data.
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 30, "total_frames": 13572, "total_tasks": 1, "total_videos": 30, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:30"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Clementppr/lerobot_pick_and_place_dataset_world_model.


See https://github.com/thuml/RLVR-World for examples for using this dataset. Citation @article{wu2025rlvr, title={RLVR-World: Training World Models with Reinforcement Learning}, author={Jialong Wu and Shaofeng Yin and Ningya Feng and Mingsheng Long}, journal={arXiv preprint arXiv:2505.13934}, year={2025}, }
