VideoWorldmodel/Evaluation
2026 Datasets
Dataset Analysis
Matrix-Game 2.0 Self-Consistency Reproduction This README explains how to reproduce the Matrix-Game 2.0 self-consistency (SC) row in Reasoning-Structured Videos: A Stratified Diagnostic Suite for Compositional Consistency in World Models. All experimental data, generated videos, and evaluation files required for this reproduction are available on the current dataset page under Files and versions: https://huggingface.co/datasets/VideoWorldmodel/Evaluation No external dataset or… See the full description on the dataset page: https://huggingface.co/datasets/VideoWorldmodel/Evaluation.
Provenance
Collected from huggingface-datasets.
downloads=0, likes=0
Related papers
$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation$ω$-EVA: Envision, Verify, and Act with Latent Interactive World ModelsA 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City SignatureA Mechanistic View on Video Generation as World Models: State and DynamicsABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics AlignmentAGEL-Comp: A Neuro-Symbolic Framework for Compositional Generalization in Interactive AgentsARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous ControlAdvancing Open-source World ModelsAgentic World Modeling: Foundations, Capabilities, Laws, and BeyondAlignUSER: Human-Aligned LLM Agents via World Models for Recommender System EvaluationAttacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World ModelsAutoregressive Diffusion World Models for Off-Policy Evaluation of LLM AgentsBRo-JEPA: Learning Modular Arithmetic in Latent SpaceBehavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement LearningBeyond Pixel Histories: World Models with Persistent 3D StateBridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment MasksCWM: Contrastive World Models for Action Feasibility Learning in Embodied Agent PipelinesCausal Object-Centric Models for Planning with Monte Carlo Tree SearchCausalDrive: Real-time Causal World Models for Autonomous DrivingChronoMedicalWorld: A Medical World Model for Learning Patient Trajectories from Longitudinal Care DataConformal Orbit-Valid Trust Horizons for Equivariant World ModelsCurrent World Models Lack a Persistent State CoreData-Asymmetric Latent Imagination and Reranking for 3D Robotic Imitation LearningDeTrack: A Benchmark and Altitude-Aware Dual World Model for Drone-embodied TrackingDeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous DrivingDiffusion Transformer World-Action Model for AV Scene PredictionDiscrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy LearningDistill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous DrivingDo World Action Models Generalize Better than VLAs? A Robustness StudyDreamDojo: A Generalist Robot World Model from Large-Scale Human VideosDreamWorld: Unified World Modeling in Video GenerationDreamX-World 1.0: A General-Purpose Interactive World ModelDreaming Of Others: Latent Teammate Modeling In World Models For Multi-Agent Reinforcement LearningDreaming Smoothly and Sample Efficiently with Gradient Penalized Latent DynamicsDriveCtrl: Conditioned Sim-to-Real Driving Video GenerationDriver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics RolloutDrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous DrivingDynaWM: Dynamics-Aware Distillation with World Model and Momentum Targets for Smooth Locomotion over Continuous StairsEcho-Memory: A Controlled Study of Memory in Action World ModelsEgoCS-400K: An Egocentric Gameplay Dataset for World ModelsEmbody4D: A Generalist Data Engine for Embodied 4D World ModelingFAR-Drive: Frame-AutoRegressive Video Generation in Closed-Loop Autonomous DrivingFrom Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World ModelsGE-Sim 2.0: A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic ManipulationGEM: Generating LiDAR World Model via Deformable MambaGeneralization of World Models under Environmental Variability for Vision-based Quadrotor NavigationGeoWorld: Geometric World ModelsHERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and GenerationHOLO-MPPI: Multi-Scenario Motion Planning via Hierarchical Policy OptimizationHarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action ModelsHi-WM: Human-in-the-World-Model for Scalable Robot Post-TrainingHow Should World Models Be Evaluated? A Decision-Making-Centric PositionIOI: Decoupling Kinematics and Physics for Interactive World ModelsInSpatio-WorldFM: An Open-Source Real-Time Generative Frame ModelIncantation: Natural Language as the Action Interface for Multi-Entity Video World ModelsIs Your Driving World Model an All-Around Player?JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement LearningLaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot ManipulationLaWM: Least Action World Models for Long-Horizon Physical Consistency from Visual ObservationsLatent State Design for World Models under Sufficiency ConstraintsLatent World Models for Automated Driving: A Unified Taxonomy, Evaluation Framework, and Open ChallengesLatent-WAM: Latent World Action Modeling for End-to-End Autonomous DrivingLearning Invariant Visual Representations for Planning with Joint-Embedding Predictive World ModelsLearning Latent Action World Models In The WildLight Interaction: Training-Free Inference Acceleration for Interactive Video World ModelsLiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World ModelsLoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)MBench: A Comprehensive Benchmark on Memory Capability for Video World ModelsMCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP EnvironmentsMIND: Benchmarking Memory Consistency and Action Control in World ModelsMODIP: Efficient Model-Based Optimization for Diffusion PoliciesMaineCoon: Pursuing A Real-Time Audio-Visual Social World ModelMaking Foresight Actionable: Repurposing Representation Alignment in World Action ModelsMask World Model: Predicting What Matters for Robust Robot Policy LearningMem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot ManipulationMind Dreamer: Untethering Imagination via Active Causal Intervention on Latent ManifoldsMobileDreamer: Generative Sketch World Model for GUI AgentMonte Carlo Pass Search: Using Trajectory Generation for 3D Counterfactual Pass Evaluation in FootballNVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle SimulationNano World Models: A Minimalist Implementation of Future Video PredictionNarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio DramaNavWAM: A Navigation World Action Model for Goal-Conditioned Visual NavigationNetwork-Efficient World Model Token StreamingNeuroHex: Highly-Efficient Hex Coordinate System for Creating World Models to Enable Adaptive AINous: A Predictive World Model for Long-Term Agent MemoryOSCAR: Omni-Embodiment Action-Conditioned World Model for RoboticsOccDirector: Language-Guided Behavior and Interaction Generation in 4D Occupancy SpaceOmni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World ModelsOmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video GenerationOn Memory: A comparison of memory mechanisms in world modelsOne Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial PerceptionOrchestrated Reality: From Role-Play to Living, Playable Game Worlds -- LLM-Driven World Simulation as a Parameterized-Action POMDPPLUME: Probabilistic Latent Unified World Modeling and Parameter Estimation for Multi-Finger ManipulationPROWL: Prioritized Regret-Driven Optimization for World Model LearningPanoWorld: A Generative Spatial World Model for Consistent Whole-House Panorama SynthesisPanoWorld: Geometry-Consistent Panoramic Video World ModelingPathWise: Planning through World Model for Automated Heuristic Design via Self-Evolving LLMsPearlVLA: Progressive Embodied Action-Plan Refinement in Latent SpacePhyGround: Benchmarking Physical Reasoning in Generative World ModelsPhyWorld: Physics-Faithful World Model for Video GenerationPhys-JEPA: Physics-Informed Latent World Models for Multivariate Time-Series ForecastingPhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World ModelsPiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop EvaluationPlanning in 8 Tokens: A Compact Discrete Tokenizer for Latent World ModelPlanning with an Ensemble of World ModelsPrisma-World: Camera-Controllable Multi-Agent Video World ModelProbabilistic Dreaming for World ModelsProbing the effectiveness of World Models for Spatial Reasoning through Test-time ScalingQuantitative Video World Model Evaluation for Geometric-ConsistencyQwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video GenerationR2-Dreamer: Redundancy-Reduced World Models without Decoders or AugmentationRAE-NWM: Navigation World Model in Dense Visual Representation SpaceRAYNOVA: Scale-Temporal Autoregressive World Modeling in Ray SpaceReactSim-Bench: Benchmarking Reactive Behavior World Model Simulation in Autonomous DrivingReactiveGWM: Steering NPC in Reactive Game World ModelsReconstruction or Semantics? What Makes a Latent Space Useful for Robotic World ModelsReference-Free Assessment of Physical Consistency in World Model-based Video GenerationRisk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous DrivingRobust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video GenerationSANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion TransformerSCAR: Self-Supervised Continuous Action Representation LearningST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D GenerationSWEET: Sparse World Modeling with Image Editing for Embodied Task ExecutionSelf-Supervised JEPA-based World Models for LiDAR Occupancy Completion and ForecastingSelf-Supervised Multi-Modal World Model with 4D Space-Time EmbeddingSolaris: Building a Multiplayer Video World Model in MinecraftSparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene RepresentationStealthy World Model Manipulation via Data PoisoningStereo World Model: Camera-Guided Stereo Video GenerationStressDream: Steering Video World Models for Robust Policy Evaluation and ImprovementSurgVista: Long-Horizon Surgical World Modeling with Plausible Instrument-Tissue DynamicsTRAP: Tail-aware Ranking Attack for World-Model PlanningTeaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State EvolutionTemporal Logic Guidance for Action-Only Diffusion Policies with World ModelsThe Trinity of Consistency as a Defining Principle for General World ModelsThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning ModelThinking with Imagination: Agentic Visual Spatial Reasoning with World SimulatorsThinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern InductionToward World Modeling of Physiological Signals with Chaos-Theoretic Balancing and Latent DynamicsTowards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future TrendsTrimming the Long-Tail of Visual World Modeling EvaluationUWM-JEPA: Predictive World Models That Imagine in Belief SpaceUniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World ModelingUnified Driving Tokens: Representation- and Geometry-Guided Discrete Tokenizer for Driving World Models and PlanningVJEPA: Variational Joint Embedding Predictive Architectures as Probabilistic World ModelsValue-guided action planning with JEPA world modelsVisual Generation Unlocks Human-Like Reasoning through Multimodal World ModelsWAM-RL: World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFTWBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model EvaluationWEAVER, Better, Faster, Longer: An Effective World Model for Robotic ManipulationWhat Makes Video World Model Latents Action-Relevant: Prediction over ReconstructionWhat-If World: A Causal Benchmark for General World Models in Embodied ScenariosWorld Action Models: A SurveyWorld Action Models: The Next Frontier in Embodied AIWorld Model Self-Distillation: Training World Models to Solve General TasksWorld Model for Robot Learning: A Comprehensive SurveyWorld Model-Enabled Causal Digital Twins for Semantic Communications in Physical AI SystemsWorld Models for Robotic Manipulation: A SurveyWorld-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied TasksWorld-R1: Reinforcing 3D Constraints for Text-to-Video GenerationWorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and PlatformWorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World ModelsWorldBench: Disambiguating Physics for Diagnostic Evaluation of World ModelslWorldCache: Content-Aware Caching for Accelerated Video World ModelsWorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World ModelsWorldFly: A World-Model-Based Vision-Language-Action Model for UAV NavigationWorldKernel: A World Model is the Coupling Kernel of Admissible Possible WorldsWorldMark: A Unified Benchmark Suite for Interactive Video World ModelsWorldOlympiad: Can Your World Model Survive a Triathlon?WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State PredictorsWorldVLM: Combining World Model Forecasting and Vision-Language ReasoningWow, wo, val! A Comprehensive Embodied World Model Evaluation Turing TestlX-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models InferenceX-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World ModelingX-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End DrivingXiaomi EV World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous DrivingYoCausal: How Far is Video Generation from World Model? A Causality PerspectivedWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Modelstable-worldmodel: A Platform for Reproducible World Modeling Research and Evaluation3D and 4D World Modeling: A Survey3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation4D Driving Scene Generation With Stereo Forcing4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation ModelsA Comprehensive Survey on World Models for Embodied AIA Survey on Future Physical World Generation for Autonomous DrivingA Survey: Learning Embodied Intelligence from Physical Simulators and World ModelsA Unified Definition of Hallucination, Or: It's the World Model, StupidAVD2: Accident Video Diffusion for Accident Video DescriptionActive Confusion Expression in Large Language Models: Leveraging World Models toward Better Social ReasoningActive Intelligence in Video Avatars via Closed-loop World ModelingAdaPower: Specializing World Foundation Models for Predictive ManipulationAdapting Vision-Language Models for Evaluating World ModelsAdapting a World Model for Trajectory Following in a 3D GameAdvancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive BenchmarksAligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AIBenchmarking World-Model LearningBeyond Generative AI: World Models for Clinical Prediction, Counterfactuals, and PlanningBootstrapping World Models from Dynamics Models in Multimodal Foundation ModelsCLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent SpaceCWM: An Open-Weights LLM for Research on Code Generation with World ModelsCausal Cartographer: From Mapping to Reasoning Over Counterfactual WorldsCausalARC: Abstract Reasoning with Causal World ModelsChronoDreamer: Action-Conditioned World Model as an Online Simulator for Robotic PlanningClone Deterministic 3D Worlds with Geometrically-Regularized World ModelsConsistent World Models via Foresight DiffusionCosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation ModelsCosmos-Transfer1Counterfactual World Models via Digital Twin-conditioned Video DiffusionCtrl-World: A Controllable Generative World Model for Robot ManipulationDMWM: Dual-Mind World Model with Long-Term ImaginationDeep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and NavigationDeep SPI: Safe Policy Improvement via World ModelsDeepVerse: 4D Autoregressive Video Generation as a World ModelDexterous World ModelsDiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene GenerationDiVE: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion TransformerDisentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement LearningDreamerV3-XP: Optimizing exploration through uncertainty estimationDreamland: Controllable World Creation with Simulator and Generative ModelsDrive&Gen: Co-Evaluating End-to-End Driving and Video Generation ModelsDynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot GeneralizationEWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World ModelsEchoWorld: Learning Motion-Aware World Models for Echocardiography Probe GuidanceEdge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and ChallengesEmu3.5: Native Multimodal Models are World LearnersEnd-to-End Driving with Online Trajectory Evaluation via BEV World ModelEnerVerse: Envisioning Embodied Future Space for Robotics ManipulationEvaluating Gemini Robotics Policies in a Veo World SimulatorEvaluating Robot Policies in a World ModelExploring the Evolution of Physics Cognition in Video Generation: A SurveyFrom 2D to 3D Cognition: A Brief Survey of General World ModelsFrom Word to World: Can Large Language Models be Implicit Text-based World Models?FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous DrivingGAIA-2: A Controllable Multi-View Generative World Model for Autonomous DrivingGEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition ControlGLAM: Global-Local Variation Awareness in Mamba-based World ModelGaussianWorld: Gaussian World Model for Streaming 3D Occupancy PredictionGenerating Multimodal Driving Scenes via Next-Scene PredictionGenerating Symbolic World Models via Test-time Scaling of Large Language ModelsGenerative Physical AI in Vision: A SurveyGenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video GenerationGigaWorld-0: World Models as Data Engine to Empower Embodied AIGrndCtrl: Grounding World Models via Self-Supervised Reward AlignmentHow Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert AssessmentHunyuan-GameCraft-2: Instruction-following Interactive Game World ModelIPR-1: Interactive Physical ReasonerKAN-Dreamer: Benchmarking Kolmogorov-Arnold Networks as Function Approximators in World ModelsKeyWorld: Key Frame Reasoning Enables Effective and Efficient World ModelsLLM-JEPA: Large Language Models Meet Joint Embedding Predictive ArchitecturesLUMOS: Language-Conditioned Imitation Learning with World ModelsLatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World GenerationLearning Abstract World Models with a Group-Structured Latent SpaceLearning Real-World Action-Video Dynamics with Heterogeneous Masked AutoregressionLearning an Adversarial World Model for Automated Curriculum Generation in MARLLearning to Generate 4D LiDAR SequencesLiDARCrafter: Dynamic 4D World Modeling from LiDAR SequencesLong-Context State-Space Video World ModelsLongLive: Real-time Interactive Long Video GenerationLongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoEManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual GuidanceMaskGWM: A Generalizable Driving World Model with Video Mask ReconstructionMatrix-Game 2.0: An Open-Source, Real-Time, and Streaming Interactive World ModelMeasuring (a Sufficient) World Model in LLMs: A Variance Decomposition FrameworkMiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous DrivingMindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous DrivingMineWorld: a Real-Time and Open-Source Interactive World Model on MinecraftMoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature ModulationMorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World SimulatorNORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference RewardsNavigation World ModelsOcc-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language ModelsOccTENS: 3D Occupancy World Model via Temporal Next-Scale PredictionOccupancy World Model for RobotsOmniGen: Unified Multimodal Sensor Generation for Autonomous DrivingOmniNWM: Omniscient Driving Navigation World ModelsOmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World ModelingOne Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided ExplorationOne Model for All Tasks: Leveraging Efficient World Models in Multi-Task PlanningOpenTwinMap: An Open-Source Digital Twin Generator for Urban Autonomous DrivingPIN-WM: Learning Physics-INformed World Models for Non-Prehensile ManipulationPlanning with Reasoning using Vision Language World ModelProTerrain: Probabilistic Physics-Informed Rough Terrain World ModelingProphetDWM: ProphetDWM: A Driving World Model for Rolling Out Future Actions and VideosRELIC: Interactive Video World Model with Long-Horizon MemoryRadarGen: Automotive Radar Point Cloud Generation from CamerasReSim: Reliable World Simulation for Autonomous DrivingRemote Sensing-Oriented World ModelRethinking Driving World Model as Synthetic Data Generator for Perception TasksRoboScape: Physics-informed Embodied World ModelSAMPO: Scale-wise Autoregression with Motion PrOmpt for generative world modelsSTORM: Search-Guided Generative World Models for Robotic ManipulationScalable Policy Evaluation with Video World ModelsScaling Up Occupancy-centric Driving Scene Generation: Dataset and MethodSceneDiffuser++: City-Scale Traffic Simulation via a Generative World ModelSekai: A Video Dataset towards World ExplorationSimple, Good, Fast: Self-Supervised World Models Free of BaggageSimuRA: Towards General Goal-Oriented Agent via Simulative Reasoning Architecture with LLM-Based World ModelSmallWorlds: Assessing Dynamics Understanding of World Models in Isolated EnvironmentsSpeech World Model: Causal State-Action Planning with Explicit Reasoning for SpeechStateSpaceDiffuser: Bringing Long Context to Diffusion World ModelsSurfer: A World Model-Based Framework for Vision-Language Robot ManipulationSynthesizing world models for bilevel planningTeraSim-World: Worldwide Safety-Critical Data Synthesis for End-to-End Autonomous DrivingTerra: Explorable Native 3D World Model with Point LatentsText2World: Benchmarking Large Language Models for Symbolic World Model GenerationThe Double Life of Code World Models: Provably Unmasking Malicious Behavior Through Execution TracesThe Safety Challenge of World Models for Embodied AI Agents: A ReviewThink Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous VehiclesThinking Ahead: Foresight Intelligence in MLLMs and World ModelsTime-Aware World Model for Adaptive Prediction and ControlToward Memory-Aided World Models: Benchmarking via Spatial ConsistencyToward Stable World Models: Measuring and Addressing World Instability in Generative EnvironmentsTowards High-Consistency Embodied World Model with Multi-View Trajectory VideosUniOcc: A Unified Benchmark for Occupancy Forecasting and Prediction in Autonomous DrivingUniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous DrivingVFMF: World Modeling by Forecasting Vision Foundation Model FeaturesVL-SAFE: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous DrivingVaViM and VaVAM: Autonomous Driving through Video Generative ModelingVideo World Models with Long-term Spatial MemoryVideoVerse: How Far is Your T2V Generator from a World Model?Vision-Centric 4D Occupancy Forecasting and Planning via Implicit Residual World ModelsVoyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene GenerationWMNav: Integrating Vision-Language Models into World Models for Object Goal NavigationWhat Does it Mean for a Neural Network to Learn a "World Model"?What Has a Foundation Model Found? Using Inductive Bias to Probe for World ModelsWhole-Body Conditioned Egocentric Video PredictionWoW: Towards a World omniscient World model Through Embodied InteractionWorld Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task PlanningWorld Models That Know When They Don't Know: Controllable Video Generation with Calibrated UncertaintyWorld Models as Reference Trajectories for Rapid Motor AdaptationWorld Models for Cognitive Agents: Transforming Edge Intelligence in Future NetworksWorld model inspired sarcasm reasoning with large language model agentsWorld-in-World: World Models in a Closed-Loop WorldWorld4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World ModelWorldLens: Full-Spectrum Evaluations of Driving World Models in Real WorldWorldModelBench: Judging Video Generation Models As World ModelsWorldPack: Compressed Memory Improves Spatial Consistency in Video World ModelingWorldVLA: Towards Autoregressive Action World ModelXray2Xray: World Model from Chest X-rays with Volumetric ContextYume-1.5: A Text-Controlled Interactive World Generation ModelZero-Splat TeleAssist: A Zero-Shot Pose Estimation Framework for Semantic TeleoperationBWArea Model: Learning World Model, Inverse Dynamics, and Policy for Controllable Language GenerationCam4DOCC: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving ApplicationsCan Language Models Serve as Text-Based World Simulators?CityBench: Evaluating the Capabilities of Large Language Model as World ModelDINO-WM: World Models on Pre-trained Visual Features enable Zero-shot PlanningDOME: Taming Diffusion Model into High-Fidelity Controllable Occupancy World ModelDiffusion for World Modeling: Visual Details Matter in AtariDoe-1: Closed-Loop Autonomous Driving with Large World ModelDream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with ImaginationDreaming of Many Worlds: Learning Contextual World Models Aids Zero-Shot GeneralizationDriveGenVLM: Real-world Video Generation for Vision Language Model based Autonomous DrivingDriving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous DrivingDrivingDiffusion: Layout-Guided multi-view driving scene video generation with latent diffusion modelDrivingWorld: Constructing World Model for Autonomous Driving via Video GPTEvaluating the World Model Implicit in a Generative ModelExploring the Interplay Between Video Generation and World Models in Autonomous Driving: A SurveyGeneralized Predictive Model for Autonomous DrivingGenie: Generative Interactive EnvironmentsHow Far is Video Generation from World Model: A Physical Law PerspectiveImagine-2-Drive: High-Fidelity World Modeling in CARLA for Autonomous VehiclesImproving Token-Based World Models with Parallel Observation PredictionLanguage Agents Meet Causality -- Bridging LLMs and Causal World ModelsLearning Latent Dynamic Robust Representations for World ModelsMaking Large Language Models into World Models with Precondition and Effect KnowledgeMotion Prompting: Controlling Video Generation with Motion TrajectoriesMulti-Task Interactive Robot Fleet Learning with Visual World ModelsMultimodal foundation world models for generalist embodied agentsOccSora: 4D Occupancy Generation Models as World Simulators for Autonomous DrivingOwl-1: Omni World Model for Consistent Long Video GenerationPanacea: Panoramic and Controllable Video Generation for Autonomous DrivingPlanning with Adaptive World Models for Autonomous DrivingProbing Multimodal LLMs as World Models for DrivingSubjectDrive: Scaling Generative Data in Autonomous Driving via Subject ControlTerra ACT-Bench: Towards Action Controllable World Models for Autonomous DrivingThink2Drive: Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous DrivingVista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityWHALE: Towards Generalizable and Scalable World Models for Embodied Decision-makingWoVoGen: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene GenerationWorld Models for Autonomous Driving: An Initial SurveyWorld Models: The Safety PerspectiveWorldSimBench: Towards Video Generation Models as World SimulatorSTORM: Efficient Stochastic Transformer based World Models for Reinforcement LearningTrafficBots: Towards World Models for Autonomous Driving Simulation and Motion PredictionTransformers are Sample Efficient World ModelsDreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical RepresentationsHierarchical Model-Based Imitation Learning for Planning in Autonomous DrivingPerceptUI: PerceptUI: LLM Agents as Human-Aligned Synthetic Users for UI/UX EvaluationMemoBench: Benchmarking World Modeling in Dynamically Changing EnvironmentsFast LeWorldModelOrca: The World is in Your MindOpenSTL: A Comprehensive Benchmark of Spatio-Temporal Predictive LearningLearning Transferable Dynamics Priors from Action to World ModelingDreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World ModelValdi: Value Diffusion World ModelsWorldDirector: Building Controllable World Simulators with Persistent Dynamic MemoryWorldOdysseyBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World ModelsFrom World Models to World Action Models: A Concise Tutorial for RoboticsOPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive ExplorationBridge-WA: Predicting Where and How the World Changes for Robotic ActionLWDrive: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous DrivingLong-term Traffic Simulation via Structured Autoregressive ModelingForgeDrive: Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous DrivingOne Video, One World: Turning Monocular Video into Physical 4D ScenesWorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World ModelsEmbodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots3D Point World Models: Point Completion Enables More Accurate Dynamics LearningRetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in RetailAGI Maze as a Benchmark Framework for World-Modeling AgentsPath Planning in Physically Viable World ModelsRoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy EvaluationIRASim: A Fine-Grained World Model for Robot ManipulationEnerVerse-AC: Envisioning Embodied Environments with Action ConditionGigaWorld-1: A Roadmap to Build World Models for Robot Policy EvaluationDynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual DynamicsLast-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing ApproachOperator-on-F complements value-equivalence: a planning-time diagnostic for latent world modelsCRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model PretrainingMask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World ModelsDSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot ManipulationMoP-JEPA: Hard-Assigned Predictor Mixtures for Stochastic JEPA World ModelsMultiplayer Interactive World Models with Representation AutoencodersAlayaWorld: Long-Horizon and Playable Video World GenerationLook Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA ModelsInfinite Worlds with Versatile InteractionsScaling Mixture-of-Experts Video Pretraining for Embodied IntelligenceV-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation ModelsXiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation ModelVersatile Behavior Diffusion for Generalized Traffic Agent SimulationLearning to drive from a world on railsVideo = World + Event StreamBadWAM: When World-Action Models Dream Right but Act WrongSelf in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied IntelligenceCollision Avoidance Detour for Multi-Agent Trajectory ForecastingFrom Pixels to States: Rethinking Interactive World Models as Game EnginesRethinking Video Generation Model for the Embodied WorldMixture of Contexts for Long Video GenerationRecurrent Autoregressive Diffusion: Global Memory Meets Local AttentionE3DGS: Unified Geometric-Photometric Equivariance for 3D Gaussian Splatting via Color-as-Geometry EmbeddingOrbis 2: A Hierarchical World Model for DrivingGeoWorldAD: Geometry World Action Model for Autonomous DrivingThinking in Video: Can Video Generators Really Reason About the Real World?Reinforcement Learning: From Algorithms To Foundation ModelsMobile Network Control with a World ModelSAGE: Subgoal-Conditioned Action Generation for Latent World Model PlanningMasked Visual Actions for Unified World ModelingMasked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RLSeerGuard: A Safety Framework for Mobile GUI Agents via World Model PredictionOpen-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied LearningApple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical IntelligenceEvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary WorldGenerative World Renderer at the Speed of PlayABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPUStreaming Multi-Agent Autoregressive Diffusion Model with World State RegistersPlayable Video GenerationPlayable Environments: Video Manipulation in Space and TimeVideoPhy: Evaluating Physical Commonsense for Video GenerationVideo Generation Models in Robotics - Applications, Research Challenges, Future DirectionsOSCBench: Benchmarking Object State Change in Text-to-Video GenerationWonder: Video World Model Done BetterWorldDiT: A Unified Diffusion Architecture for World and Action ModelingTemporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive ControlVisualPatchWorld: Code World Models as Latent Structured Representations for PlanningACE-Data-0: Human-Centric Ambient Capture as Embodied Data EngineStatePlay: State-Aware Game World Models for Mechanics-Consistent GenerationCuriosity-Driven Exploration by Self-Supervised PredictionPlanning from Pixels using Inverse Dynamics ModelsSG-WAM: Self-Guided World Modeling in Geometry-Aware Policy SpaceHelloWorld: Enabling Socially Interactive Characters in Video World ModelsMiniWorld: Democratizing the Training of Video World Models from ScratchDyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment ManipulationWorldClaw: Agentic 3D Open-World Generation at Scale
Resources
Tags
region:usworld-modelvideo-generationbenchmarkreproducibility