qsun2001/world_model
2026 Datasets
Dataset Analysis
Hugging Face dataset: qsun2001/world_model
Provenance
Collected from huggingface-datasets.
downloads=3152, likes=0
Related papers
$μ_0$: A Scalable 3D Interaction-Trace World Model$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation$ω$-EVA: Envision, Verify, and Act with Latent Interactive World Models3D-Belief: Embodied Belief Inference via Generative 3D World ModelingA 3D Isovist World Model -- Revealing a City's Unseen Geometry and Its Emergent Cross-City SignatureA Close Look At World Model Recovery In Supervised Fine-Tuned LLM PlannersA geometric relation of the error introduced by sampling a language model's output distribution to its internal stateAGEL-Comp: A Neuro-Symbolic Framework for Compositional Generalization in Interactive AgentsAGWM: Affordance-Grounded World Models for Environments with Compositional PrerequisitesAR Forcing: Towards Long-Horizon Robot Navigation World ModelARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous ControlARIS: Agentic and Relationship Intelligence System for Social RobotsActWorld: From Explorable to Interactive World Model via Action-Aware MemoryActive Inference as the Test-Time Scaling Law for Physical AI AgentsActive Inference: A method for Phenotyping Agency in AI systems?AdaReP:Adaptive Re-Planning under Model Mismatch for Neural World-Model Predictive ControlAdaptiveLoad: Towards Efficient Video Diffusion Transformer TrainingAffectVerse: Emotional World Models for Multimodal Affective ComputingAffective Music Recommendation: A Rollout-Based World Model for Offline Preference OptimizationAgentic World Modeling: Foundations, Capabilities, Laws, and BeyondAgentifying Patient Dynamics within LLMs through Interacting with Clinical World ModelAttacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World ModelsAutonomous Video Generation with Counterfactual Controllability for Self-Evolving World ModelsAutoregressive Diffusion World Models for Off-Policy Evaluation of LLM AgentsBRo-JEPA: Learning Modular Arithmetic in Latent SpaceBadDreamer: Transferable Backdoor Attacks against Video World Models for Autonomous DrivingBehavior-Invariant Task Representation Learning with Transformer-based World Models for Offline Meta-Reinforcement LearningBiWM: Advancing Open-Source Interactive Video World Models with Bidirectional AutoregressionBusiness World ModelCLAW: Learning Continuous Latent Action World Models via Adversarial Latent RegularizationCOMAP: Co-Evolving World Models and Agent Policies for LLM AgentsCan In-Context Learning Support Intrinsic Curiosity?Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video GenerationCausal Object-Centric Models for Planning with Monte Carlo Tree SearchCausal Reward World Models: Zero-shot Reward Design for Automated Skill GenerationCausal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World ModelsCausalDrive: Real-time Causal World Models for Autonomous DrivingCertified World Models: Predictability Across Configuration, Horizon, and ResolutionChreode: A Cell World Model for One-Step Temporal Dynamics and Perturbation PredictionChronoMedicalWorld: A Medical World Model for Learning Patient Trajectories from Longitudinal Care DataClosing the Motion Execution Gap: From Semantic Motion Task Constraints to Kinematic ControlCoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous DrivingCoding Agent Is Good As World SimulatorConformal Orbit-Valid Trust Horizons for Equivariant World ModelsCurrent World Models Lack a Persistent State CoreDREAM-Chunk: Reactive Action Chunking with Latent World ModelData-Asymmetric Latent Imagination and Reranking for 3D Robotic Imitation LearningDeTrack: A Benchmark and Altitude-Aware Dual World Model for Drone-embodied TrackingDecMem: Towards Minute-Long Consistent World Generation with Decoupled MemoryDeepSight: Long-Horizon World Modeling via Latent States Prediction for End-to-End Autonomous DrivingDeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from VideosDelta Forcing: Trust Region Steering for Interactive Autoregressive Video GenerationDemo-JEPA: Joint-Embedding Predictive Architecture for One-shot Cross-Embodiment ImitationDexFuture: Hierarchical Future-State Visuomotor Targeting for Bimanual Dexterous Tool UseDiLA: Disentangled Latent Action World ModelsDiffusion Transformer World-Action Model for AV Scene PredictionDisCo: World Models with Discrete Camera Motion ControlDiscrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy LearningDistill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous DrivingDistilling Game Code World Model Generation into Lightweight Large Language ModelsDivide and Conquer: Decoupled Representation Alignment for Multimodal World ModelsDo multimodal models imagine electric sheep?Dream-MPC: Gradient-Based Model Predictive Control with Latent ImaginationDreamX-World 1.0: A General-Purpose Interactive World ModelDreaming Of Others: Latent Teammate Modeling In World Models For Multi-Agent Reinforcement LearningDreaming Smoothly and Sample Efficiently with Gradient Penalized Latent DynamicsDrift-Resistant Navigation World Model with Anchored Epipolar GuidanceDriver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics RolloutDual-Channel Grounded World Modeling (DCGWM): Structural Prevention of Objective Interference Collapse via Heterogeneous External Grounding with Inward-Only Gradient FlowDynaWM: Dynamics-Aware Distillation with World Model and Momentum Targets for Smooth Locomotion over Continuous StairsDynoSLAM: Dynamic SLAM with Generative Graph Neural Networks for Real-World Social NavigationEA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action FieldsEV-WM: Event-Verified World Models for Long-Horizon Robotic ManipulationEcho-Memory: A Controlled Study of Memory in Action World ModelsEfficient Agentic Reasoning Through Self-Regulated Simulative PlanningEgoCS-400K: An Egocentric Gameplay Dataset for World ModelsEgoExo-WM: Unlocking Exo Video for Ego World ModelsEmbodied Multi-Agent Coordination by Aligning World Models Through DialogueEmbody4D: A Generalist Data Engine for Embodied 4D World ModelingEmergent Semantic Representations in World Models through Physical Interaction without Linguistic SupervisionEmotion-Conditioned Short-Horizon Human Pose Forecasting with a Lightweight Predictive World ModelEponaV2: Driving World Model with Comprehensive Future ReasoningExact equivariance, kept through training, buys zero-shot generalisation across the symmetry groupExecutable World Models for ARC-AGI-3 in the Era of Coding AgentsExplainably Safe Reinforcement LearningFeat2Go: Visual Feature-Grounded Value Estimation for Embodied Reinforcement LearningFeedback World Model Enables Precise Guidance of Diffusion PolicyFlowMPC: Improving Flow Matching policies with World ModelsFlowMo-WM: A World Model with Object Momentum and Hidden Ambient DriftFlyMirage: A Fully Automated Generation Pipeline for Diverse and Scalable UAV Flight Data via Generative World ModelFlying by Inference: Active Inference World Models for Adaptive UAV SwarmsForesight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model LatentsFrom Pixels to Concepts: Growing Rich 3D Semantic Scene Graph Forests utilizing Foundation ModelsFrom Zero to Hero: Training-Free Custom Concept Spawning in World ModelsFuture Dynamic 3D Reconstruction: A 3D World Model with Disentangled Ego-MotionGE-Sim 2.0: A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic ManipulationGEM-4D: Geometry-Enhanced Video World Models for Robot ManipulationGEM: Gaussian Evolution Model for Occupancy Forecasting and Motion PlanningGEM: Generating LiDAR World Model via Deformable MambaGamma-World: Generative Multi-Agent World Modeling Beyond Two PlayersGaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic ManipulationGeneralization of World Models under Environmental Variability for Vision-based Quadrotor NavigationGeoSem-WAM: Geometry- and Semantic-Aware World Action ModelsGeoStream: Toward Precise Camera Controlled Streaming Video GenerationGroupToM-Bench: Benchmarking Group Theory of Mind and Nonlinear Social Emergence in MLLMsHEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World ModelsHERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and GenerationHOLO-MPPI: Multi-Scenario Motion Planning via Hierarchical Policy OptimizationHaM-World: Soft-Hamiltonian World Models with Selective Memory for PlanningHarmoWAM: Harmonizing Generalizable and Precise Manipulation via Adaptive World Action ModelsHi-WM: Human-in-the-World-Model for Scalable Robot Post-TrainingHolo-World: Unified Camera, Object and Weather Control for Video World ModelHorizonDrive: Self-Corrective Autoregressive World Model for Long-horizon Driving SimulationHow Mobile World Model Guides GUI Agents?How Should World Models Be Evaluated? A Decision-Making-Centric PositionIDOL: Inverse-Dynamics-Guided Future Prediction for End-to-End Autonomous DrivingIFPV: An Integrated Multi-Agent Framework for Generative Operational Planning and High-Fidelity Plan VerificationIMWM: Intuition Models Complement World Models for Latent PlanningIOI: Decoupling Kinematics and Physics for Interactive World ModelsIdentifiability Without Gaussianity: Symbolic World Models and Near-Infinite Temporal ConsistencyImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?Imagine to Ensure Safety in Hierarchical Reinforcement LearningImitation from Heterogeneous Demonstrations using Grounded Latent-Action World ModelsImperfect World Models are ExploitableIncantation: Natural Language as the Action Interface for Multi-Entity Video World ModelsInference-time Policy Steering via Vision and TouchIntercepting the Future: Latent-Space Predictive World Model for Dynamic VLA ManipulationInverting the Bellman Equation: From $Q$-Values to World ModelsIs Your Driving World Model an All-Around Player?JEDI: Joint Embedding Diffusion World Model for Online Model-Based Reinforcement LearningKairos: A Native World Model Stack for Physical AILASER: Learning Active Sensing for Continuum Field ReconstructionLEIA: Learned Environment for Interactive Architected MaterialsLVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving ModelLaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video GenerationLaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot ManipulationLaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot PoliciesLaWM: Least Action World Models for Long-Horizon Physical Consistency from Visual ObservationsLatent Geometry Beyond Search: Amortizing Planning in World ModelsLatent Spatial Memory for Video World ModelsLatent State Design for World Models under Sufficiency ConstraintsLearning Action-Conditional and Object-Centric Gaussian Splatting World Models for Rigid ObjectsLearning Multi-Timescale Abstractions for Hierarchical Combinatorial PlanningLearning Visual Feature-Based World Models via Residual Latent ActionLearning a Particle Dynamics Model with Real-world VideosLifting Embodied World Models for Planning and ControlLight Interaction: Training-Free Inference Acceleration for Interactive Video World ModelsLoViF 2026 The First Challenge on Holistic Quality Assessment for 4D World Model (PhyScore)MAD: Mapping-Aware World Models for Agile Quadrotor FlightMBench: A Comprehensive Benchmark on Memory Capability for Video World ModelsMCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP EnvironmentsMODIP: Efficient Model-Based Optimization for Diffusion PoliciesMaineCoon: Pursuing A Real-Time Audio-Visual Social World ModelMaking Foresight Actionable: Repurposing Representation Alignment in World Action ModelsMask World Model: Predicting What Matters for Robust Robot Policy LearningMem-World: Memory-Augmented Action-Conditioned World Models for Persistent Robot ManipulationMemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action ModelsMetaWorld: Scaling Multi-Agent Video World Model from Single-view Video DataMind Dreamer: Untethering Imagination via Active Causal Intervention on Latent ManifoldsMoVerse: Real-Time Video World Modeling with Panoramic Gaussian ScaffoldMonte Carlo Pass Search: Using Trajectory Generation for 3D Counterfactual Pass Evaluation in FootballMotionWAM: Towards Foundation World Action Models for Real-Time Humanoid Loco-ManipulationMotuBrain: An Advanced World Action Model for Robot ControlNVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle SimulationNano World Models: A Minimalist Implementation of Future Video PredictionNarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio DramaNavWAM: A Navigation World Action Model for Goal-Conditioned Visual NavigationNavWM: A Unified Navigation World Model for Foresight-Driven PlanningNetwork-Efficient World Model Token StreamingNext Forcing: Causal World Modeling with Multi-Chunk PredictionNous: A Predictive World Model for Long-Term Agent MemoryOSCAR: Omni-Embodiment Action-Conditioned World Model for RoboticsOccDirector: Language-Guided Behavior and Interaction Generation in 4D Occupancy SpaceOmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video GenerationOne Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial PerceptionOne Lens, Many Worlds : A Capability-Typed Interface for World-Model InterpretabilityOne Transit Is All You Need: Detecting Exoplanets Through Learned Stellar Behaviour with EXOVEILOptiWorld: Optimal Control for Video World Generation under Physical ConstraintsOrbiSim: World Models as Differentiable Physics Engines for Embodied IntelligencePEACE: A Planner-Executor Agent with Constraint Enforcement for UAVsPH-Dreamer: A Physics-Driven World Model via Port-Hamiltonian Generative DynamicsPLUME: Probabilistic Latent Unified World Modeling and Parameter Estimation for Multi-Finger ManipulationPRISM: PRior-guided Imagination Sampling in world ModelsPROWL: Prioritized Regret-Driven Optimization for World Model LearningPanoWorld: A Generative Spatial World Model for Consistent Whole-House Panorama SynthesisPanoWorld: Geometry-Consistent Panoramic Video World ModelingPatchWorld: Gradient-Free Optimization of Executable World ModelsPearlVLA: Progressive Embodied Action-Plan Refinement in Latent SpacePhyGround: Benchmarking Physical Reasoning in Generative World ModelsPhyWorld: Physics-Faithful World Model for Video GenerationPhys-JEPA: Physics-Informed Latent World Models for Multivariate Time-Series ForecastingPhysical Object Understanding with a Physically Controllable World ModelPhysically Native World Models: A Hamiltonian Perspective on Generative World ModelingPhysics-Aware Sparse Learning and Selective Online Adaptation for Euler-Lagrange Robot DynamicsPhysics-IQ VerifiedPiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop EvaluationPixels to Proofs: Probabilistically-Safe Latent World Model Control via Parallel Conformal Robust MPCPre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model RolloutsPredictive but Not Plannable: RC-aux for Latent World ModelsPriorZero: Bridging Language Priors and World Models for Decision MakingPrisma-World: Camera-Controllable Multi-Agent Video World ModelProDrive: Proactive Planning for Autonomous Driving via Ego-Environment Co-EvolutionProPlay: Procedural World Models for Self-Evolving LLM AgentsProbing the Impact of Scale on Data-Efficient, Generalist Transformer World Models for AtariQuantitative Video World Model Evaluation for Geometric-ConsistencyQuantum Cinema: An Interactive Cinematic Exploration of Quantum Computing Hardware via Generative World ModelsQwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video GenerationReactSim-Bench: Benchmarking Reactive Behavior World Model Simulation in Autonomous DrivingReactiveGWM: Steering NPC in Reactive Game World ModelsReason--Imagine--Act: Closed-Loop LLM Decision Making with World Models for Autonomous DrivingReconstruction or Semantics? What Makes a Latent Space Useful for Robotic World ModelsReference-Free Assessment of Physical Consistency in World Model-based Video GenerationReflectiChain: Epistemic Grounding in LLM-Driven World Models for Supply Chain ResilienceReinforcing VLAs in Task-Agnostic World ModelsRender, Don't Decode: Weight-Space World Models with Latent Structural DisentanglementRiding the Shifting Potential: When Reactive Control Suffices for Multi-Goal BehaviorRoboDream: Compositional World Models for Scalable Robot Data SynthesisRoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic ManipulationSANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion TransformerSC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video GenerationSCAR: Self-Supervised Continuous Action Representation LearningSIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World ModelST-Gen4D: Embedding 4D Spatiotemporal Cognition into World Model for 4D GenerationSURGE: Approximation and Training Free Particle Filter for Diffusion SurrogateSWAP: Symmetric Equivariant World-Model for Agile Robot ParkourSWEET: Sparse World Modeling with Image Editing for Embodied Task ExecutionSWoMo: Neuro-Symbolic World Model for Cataract Surgery SimulationSafeDojo: Safe Reinforcement Learning for VLA via Interactive World ModelSafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead ReasoningScaling World-Model Reinforcement Learning Through Diffusion Policy OptimizationSee Tomorrow, Act Today: Foresight-Driven Autonomous DrivingSelf-Evolving Cognitive Framework via Causal World Modeling for Embodied Scientific IntelligenceSelf-supervised Hierarchical Visual Reasoning with World ModelSensorimotor World Models: Perception for Action via Inverse DynamicsShadow-Loom: Causal Reasoning over Graphical World Models of NarrativesSigned Compression Progress on a Sealed Audit is Goodhart-ResistantSimulating clinical interventions with a generative multimodal model of human physiologySkyJEPA: Learning Long-Horizon World Models for Zero-Shot Sim-to-Real Control of QuadrotorsSlot-MPC: Goal-Conditioned Model Predictive Control with Object-Centric RepresentationsSocial World Model for Lifelong Social IntelligenceSparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene RepresentationStealthy World Model Manipulation via Data PoisoningStressDream: Steering Video World Models for Robust Policy Evaluation and ImprovementStructure Abstraction and Generalization in a Hippocampal-Entorhinal Inspired World ModelSub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World ModelsSurgVista: Long-Horizon Surgical World Modeling with Plausible Instrument-Tissue DynamicsSword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-TrainingTRAP: Tail-aware Ranking Attack for World-Model PlanningTacForeSight: Force-Guided Tactile World Model for Contact-Rich ManipulationTargeting World Models to Compromise Robot Learning PipelinesTeaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State EvolutionTemporal Logic Guidance for Action-Only Diffusion Policies with World ModelsThinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern InductionThoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement PlanningThree-in-One World Model: Energy-Based Consistency, Prediction, and Counterfactual Inference for Marketing InterventionToward AI That Understands Self and Others: A World-Model Theory of Cognitive Diversity and AlignmentToward Compiler World Models: Learning Latent Dynamics for Efficient Tensor Program SearchToward Safe Autonomous Robotic Endovascular Interventions using World ModelsToward World Modeling of Physiological Signals with Chaos-Theoretic Balancing and Latent DynamicsTowards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future TrendsTransformers Linearly Represent Highly Structured World ModelsTrimming the Long-Tail of Visual World Modeling EvaluationTurning Video Models into Generalist Robot PoliciesUSS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics LearningUWM-JEPA: Predictive World Models That Imagine in Belief SpaceUniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World ModelingUnified 3D Scene Understanding Through Physical World ModelingUnified 4D World Action Modeling from Video Priors with Asynchronous DenoisingUnified Driving Tokens: Representation- and Geometry-Guided Discrete Tokenizer for Driving World Models and PlanningUnifying Object-Centric World Models and Diffusion Policy: A Hierarchical Framework for Multi-Stage Robotic TasksUniviewVLA: A Unified Multiview Vision-Language-Action Model with World ModelingVANDERER: Map-Free Exploration using Future-Aware and Visual-Curiosity-Guided Diffusion PolicyVegSim: A Geospatial World Model for Scenario-Conditioned Vegetation SimulationVideo Generation with Predictive LatentsWAM-RL: World-Action Model Reinforcement Learning with Reconstruction Rewards and Online Video SFTWBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model EvaluationWEAVER, Better, Faster, Longer: An Effective World Model for Robotic ManipulationWh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation DataWhat Makes Video World Model Latents Action-Relevant: Prediction over ReconstructionWhat You Think is What You See: Driving Exploration in VLM Agents via Visual-Linguistic CuriosityWhat-If World: A Causal Benchmark for General World Models in Embodied ScenariosWhen Does LeJEPA Learn a World Model?Why Conclusions Diverge from the Same Observations: Formalizing World-Model Non-Identifiability via an InferenceWords as Difference Makers: How Large Language Models Determine Causal Structure in TextWorld Action Models: A SurveyWorld Action Models: The Next Frontier in Embodied AIWorld Model Self-Distillation: Training World Models to Solve General TasksWorld Model for Robot Learning: A Comprehensive SurveyWorld Model-Enabled Causal Digital Twins for Semantic Communications in Physical AI SystemsWorld Models for Robotic Manipulation: A SurveyWorld Models in Pieces: Structural Certification for General AgentsWorld Pilot: Steering Vision-Language-Action Models with World-Action PriorsWorld-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied TasksWorld-Task Factorization for Robot LearningWorld2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial ReasoningWorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and PlatformWorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World ModelsWorldFly: A World-Model-Based Vision-Language-Action Model for UAV NavigationWorldKernel: A World Model is the Coupling Kernel of Admissible Possible WorldsWorldMark: A Unified Benchmark Suite for Interactive Video World ModelsWorldOlympiad: Can Your World Model Survive a Triathlon?WorldString: Actionable World RepresentationX-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models InferenceX-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World ModelingXiaomi Auto World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous DrivingYoCausal: How Far is Video Generation from World Model? A Causality PerspectivedWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World ModelminWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Modelsstable-worldmodel: A Platform for Reproducible World Modeling Research and EvaluationPaper - Voyager: An Open-Ended Embodied Agent with Large Language ModelsA Mechanistic View on Video Generation as World Models: State and DynamicsA formal theory on problem space as a semantic world model in systems engineeringABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics AlignmentAcceRL: A Distributed Asynchronous Reinforcement Learning and World Model Framework for Vision-Language-Action ModelsAction Shapley: A Training Data Selection Metric for World Model in Reinforcement LearningAdvancing Open-source World ModelsAgent World Model: Infinity Synthetic Environments for Agentic Reinforcement LearningAlignUSER: Human-Aligned LLM Agents via World Models for Recommender System EvaluationAligning Agentic World Models via Knowledgeable Experience LearningAn Efficient and Multi-Modal Navigation System with One-Step World ModelBeyond Pixel Histories: World Models with Persistent 3D StateBoltzmann-GPT: Bridging Energy-Based World Models and Language GenerationBridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment MasksBridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion RepresentationCWM: Contrastive World Models for Action Feasibility Learning in Embodied Agent PipelinesCausal World Modeling for Robot ControlCross-View World ModelsCurrent Agents Fail to Leverage World Model as Tool for ForesightDCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video GenerationDescribe-Then-Act: Proactive Agent Steering via Distilled Language-Action World ModelsDo World Action Models Generalize Better than VLAs? A Robustness StudyDreamDojo: A Generalist Robot World Model from Large-Scale Human VideosDreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World ModelsDreamSAC: Learning Hamiltonian World Models via Symmetry ExplorationDreamWorld: Unified World Modeling in Video GenerationDreamerAD: Efficient Reinforcement Learning via Latent World Model for Autonomous DrivingDrive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End DrivingDriveCtrl: Conditioned Sim-to-Real Driving Video GenerationDrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous DrivingDynVLA: Learning World Dynamics for Action Reasoning in Autonomous DrivingEVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics RewardsEnhance Sample Efficiency and Robustness of End-to-end Urban Autonomous Driving via Semantic Masked World ModelExplicit World Models for Reliable Human-Robot CollaborationFactored Latent Action World ModelsFlow Equivariant World Models: Memory for Partially Observed Dynamic EnvironmentsFoundation World Models for Agents that Learn, Verify, and Adapt Reliably Beyond Static EnvironmentsFrom Generative Engines to Actionable Simulators: The Imperative of Physical Grounding in World ModelsFrom Observations to Events: Event-Aware World Model for Reinforcement LearningFrom Part to Whole: 3D Generative World Model with an Adaptive Structural HierarchyGeoWorld: Geometric World ModelsGigaBrain-0.5M: a VLA That Learns From World Model-Based Reinforcement LearningHand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand GesturesImagine-then-Plan: Agent Learning from Adaptive Lookahead with World ModelsInSpatio-WorldFM: An Open-Source Real-Time Generative Frame ModelInference-time Physics Alignment of Video Generative Models with Latent World ModelsInterpreting Physics in Video World ModelsKinematics-Aware Latent World Models for Data-Efficient Autonomous DrivingLIVE: Long-horizon Interactive Video World ModelingLatent World Models for Automated Driving: A Unified Taxonomy, Evaluation Framework, and Open ChallengesLatent-WAM: Latent World Action Modeling for End-to-End Autonomous DrivingLearning Invariant Visual Representations for Planning with Joint-Embedding Predictive World ModelsLearning Latent Action World Models In The WildLiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World ModelsMAD: Motion Appearance Decoupling for efficient Driving World ModelsMIND: Benchmarking Memory Consistency and Action Control in World ModelsMMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and GenerationMVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic ManipulationMWM: Mobile World Models for Action-Conditioned Consistent PredictionMetaOthello: A Controlled Study of Multiple World Models in TransformersMetaWorld: Skill Transfer and Composition in a Hierarchical World Model for Grounding High-Level InstructionsMobileDreamer: Generative Sketch World Model for GUI AgentModel Predictive Control with Differentiable World Models for Offline Reinforcement LearningMosaicMem: Hybrid Spatial Memory for Controllable Video World ModelsNeoVerse: Enhancing 4D World Model with in-the-wild Monocular VideosNeuroHex: Highly-Efficient Hex Coordinate System for Creating World Models to Enable Adaptive AIOCCVAR: Scalable 4D Occupancy Prediction via Next-Scale PredictionObject-Centric World Models Meet Monte Carlo Tree SearchObject-Centric World Models for Causality-Aware Reinforcement LearningOlaf-World: Orienting Latent Actions for Video World ModelingOmni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World ModelsOn Memory: A comparison of memory mechanisms in world modelsOut of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World ModelsPathWise: Planning through World Model for Automated Heuristic Design via Self-Evolving LLMsPersistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement LearningPhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World ModelsPlanning in 8 Tokens: A Compact Discrete Tokenizer for Latent World ModelPlanning with an Ensemble of World ModelsPointWorld: Scaling 3D World Models for In-The-Wild Robotic ManipulationProbabilistic Dreaming for World ModelsProbing the effectiveness of World Models for Spatial Reasoning through Test-time ScalingPuzzle it Out: Local-to-Global World Model for Offline Multi-Agent Reinforcement LearningR2-Dreamer: Redundancy-Reduced World Models without Decoders or AugmentationRAE-NWM: Navigation World Model in Dense Visual Representation SpaceRAYNOVA: Scale-Temporal Autoregressive World Modeling in Ray SpaceReWorld: Multi-Dimensional Reward Modeling for Embodied World ModelsReal Real-Time Long Video Generation ModelResWM: Residual-Action World Model for Visual RLResWorld: Temporal Residual World Model for End-to-End Autonomous DrivingRisk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous DrivingSay, Dream, and Act: Learning Video World Models for Instruction-Driven Robot ManipulationScaling World Model for Hierarchical Manipulation PoliciesSelf-Improving World Modelling with Latent ActionsSelf-Supervised JEPA-based World Models for LiDAR Occupancy Completion and ForecastingSelf-Supervised Multi-Modal World Model with 4D Space-Time EmbeddingSemantic Belief-State World Model for 3D Human Motion PredictionShareVerse: Multi-Agent Consistent Video Generation for Shared World ModelingSimulation Distillation: Pretraining World Models in Simulation for Rapid Real-World AdaptationSolaris: Building a Multiplayer Video World Model in MinecraftStereo World Model: Camera-Guided Stereo Video GenerationThe Trinity of Consistency as a Defining Principle for General World ModelsThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning ModelToward Physically Consistent Driving Video World Models under Challenging TrajectoriesUCM: Unifying Camera Control and Memory with Time-aware Positional Encoding Warping for World ModelsUniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous DrivingUniFuture: A 4D Driving World Model for Future Generation and PerceptionVJEPA: Variational Joint Embedding Predictive Architectures as Probabilistic World ModelsVLA-JEPA: Enhancing Vision-Language-Action Model with Latent World ModelVLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World ModelVLM-DEWM: Dynamic External World Model for Verifiable and Resilient Vision-Language Planning in ManufacturingValue-guided action planning with JEPA world modelsVega: Learning to Drive with Natural Language InstructionsVerseCrafter: Dynamic Realistic Video World Model with 4D Geometric ControlVisual Generation Unlocks Human-Like Reasoning through Multimodal World ModelsWalk through Paintings: Egocentric World Models from Internet PriorsWhat Drives Success in Physical Planning with Joint-Embedding Predictive World Models?When World Models Dream Wrong: Physical-Conditioned Adversarial Attacks against World ModelsWorld Action Models are Zero-shot PoliciesWorld Properties without World Models: Recovering Spatial and Temporal Structure from Co-occurrence Statistics in Static Word EmbeddingsWorld-VLA-Loop: Closed-Loop Learning of Video World Model and VLA PolicyWorld2Act: Latent Action Post-Training via Skill-Compositional World ModelsWorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World ModelsWorldBench: Disambiguating Physics for Diagnostic Evaluation of World ModelslWorldCache: Accelerating World Models for Free via Heterogeneous Token CachingWorldCache: Content-Aware Caching for Accelerated Video World ModelsWorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous DrivingWorldVLM: Combining World Model Forecasting and Vision-Language ReasoningWow, wo, val! A Comprehensive Embodied World Model Evaluation Turing TestlX-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End DrivingXiaomi EV World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving3D and 4D World Modeling: A Survey3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model4DWorldBench: A Comprehensive Evaluation Framework for 3D/4D World Generation ModelsA "Good" Regulator May Provide a World Model for Intelligent SystemsA Comprehensive Survey on World Models for Embodied AIA Recipe for Efficient Sim-to-Real Transfer in Manipulation with Online Imitation-Pretrained World ModelsA Step Toward World Models: A Survey on Robotic ManipulationA Survey of World Models for Autonomous DrivingA Survey on Future Physical World Generation for Autonomous DrivingA Survey on World Models Grounded in Acoustic Physical InformationA Survey: Learning Embodied Intelligence from Physical Simulators and World ModelsA Unified Definition of Hallucination, Or: It's the World Model, StupidAD-L-JEPA: Self-Supervised Spatial World Models with Joint Embedding Predictive Architecture for Autonomous Driving with LiDAR DataAD-R1: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World ModelsAVD2: Accident Video Diffusion for Accident Video DescriptionAct2Goal: From World Model To General Goal-conditioned PolicyActive Confusion Expression in Large Language Models: Leveraging World Models toward Better Social ReasoningActive Intelligence in Video Avatars via Closed-loop World ModelingAdaPower: Specializing World Foundation Models for Predictive ManipulationAdaWM: Adaptive World Model based Planning for Autonomous DrivingAdapting Vision-Language Models for Evaluating World ModelsAdapting a World Model for Trajectory Following in a 3D GameAdvancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive BenchmarksAerial World Model for Long-horizon Visual Generation and Navigation in 3D SpaceAether: Geometric-Aware Unified World ModelingAgentic World Modeling for 6G: Near-Real-Time Generative State-Space ReasoningAligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AIAstra: General Interactive World Model with Autoregressive DenoisingAstraNav-World: World Model for Foresight Control and ConsistencyAudio-Visual World Models: Towards Multisensory Imagination in Sight and SoundBack to the Features: DINO as a Foundation for Video World ModelsBenchmarking World-Model LearningBetter World Models Can Lead to Better Post-Training PerformanceBeyond Generative AI: World Models for Clinical Prediction, Counterfactuals, and PlanningBiTAgent: A Task-Aware Modular Framework for Bidirectional Coupling between Multimodal Large Language Models and World ModelsBootstrapping World Models from Dynamics Models in Multimodal Foundation ModelsBridging the Gap Between Multimodal Foundation Models and World ModelsCLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent SpaceCVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous DrivingCWM: An Open-Weights LLM for Research on Code Generation with World ModelsCausal Cartographer: From Mapping to Reasoning Over Counterfactual WorldsCausalARC: Abstract Reasoning with Causal World ModelsChronoDreamer: Action-Conditioned World Model as an Online Simulator for Robotic PlanningClone Deterministic 3D Worlds with Geometrically-Regularized World ModelsClosing the Train-Test Gap in World Models for Gradient-Based PlanningCo-Evolving Latent Action World ModelsCoEx -- Co-evolving World-model and ExplorationCoIRL-AD: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous DrivingCoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action ModelsCode World Models for General Game PlayingConsistent World Models via Foresight DiffusionContext and Diversity Matter: The Emergence of In-Context Learning in World ModelsContinual Reinforcement Learning by Planning with Online World ModelsCorrectAD: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous DrivingCounterfactual World Models via Digital Twin-conditioned Video DiffusionCritiques of World ModelsCtrl-World: A Controllable Generative World Model for Robot ManipulationDC-MPC: Discrete Codebook World Models for Continuous ControlDINO-Foresight: Looking into the Future with DINODIO: Decomposable Implicit 4D Occupancy-Flow World ModelDMWM: Dual-Mind World Model with Long-Term ImaginationDR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent CollaborationDREAMer-VXS: A Latent World Model for Sample-Efficient AGV Exploration in Stochastic, Unobserved EnvironmentsDSG-World: Learning a 3D Gaussian World Model from Dual State VideosDeductive Chain-of-Thought Augmented Socially-aware Robot Navigation World ModelDeep Active Inference with Diffusion Policy and Multiple Timescale World Model for Real-World Exploration and NavigationDeep SPI: Safe Policy Improvement via World ModelsDeepVerse: 4D Autoregressive Video Generation as a World ModelDesign and Optimization of Reinforcement Learning-Based Agents in Text-Based GamesDeterministic World Models for Verification of Closed-loop Vision-based SystemsDexterous World ModelsDiVE: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion TransformerDiWA: Diffusion Policy Adaptation with World ModelsDisentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement LearningDoes End-to-End Autonomous Driving Really Need Perception Tasks?Dream to Drive with Predictive Individual World ModelDream to Drive: Model-Based Vehicle Control Using Analytic World ModelsDreamerV3-XP: Optimizing exploration through uncertainty estimationDriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion AlignmentDrive&Gen: Co-Evaluating End-to-End Driving and Video Generation ModelsDriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene RepresentationDriveLaW: Unifying Planning and Video Generation in a Latent Driving WorldDriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous DrivingDrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive TransformersDual-Mind World Models: A General Framework for Learning in Dynamic Wireless NetworksDual-Stream Diffusion for World-Model Augmented Vision-Language-Action ModelDyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile ManipulationDyn-O: Building Structured World Models with Object-Centric RepresentationsDyna-Think: Synergizing Reasoning, Acting, and World Model Simulation in AI AgentsDynamic Sparsity: Challenging Common Sparsity Assumptions for Learning World Models in Robotic Reinforcement Learning BenchmarksDynamicCity: Large-Scale LiDAR Generation from Dynamic ScenesDynamics-Aligned Latent Imagination in Contextual World Models for Zero-Shot GeneralizationEWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World ModelsEchoWorld: Learning Motion-Aware World Models for Echocardiography Probe GuidanceEdge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and ChallengesEfficient Generation of Diverse Cooperative Agents with World ModelsEgo-Vision World Model for Humanoid Contact PlanningEmbodied AI Agents: Modeling the WorldEmbodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World ModelEmbodied World Models Emerge from Navigational Task in Open-Ended EnvironmentsEmu3.5: Native Multimodal Models are World LearnersEnd-to-End Driving with Online Trajectory Evaluation via BEV World ModelEnhancing Physical Consistency in Lightweight World ModelsEpona: Autoregressive Diffusion World Model for Autonomous DrivingEvaluating Gemini Robotics Policies in a Veo World SimulatorEvaluating Robot Policies in a World ModelEvoAgent: Agent Autonomous Evolution with Continual World Model for Long-Horizon TasksFASTopoWM: Fast-Slow Lane Segment Topology Reasoning with Latent World ModelsFLARE: Robot Learning with Implicit World ModelingFOUNDER: Grounding Foundation Models in World Models for Open-Ended Embodied Decision MakingFUTURIST: Advancing Semantic Future Prediction through Multimodal Visual Sequence TransformersFantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D PredictionFieldSeer I: Physics-Guided World Models for Long-Horizon Electromagnetic Dynamics under Partial ObservabilityFlowDreamer: A RGB-D World Model with Flow-based Motion Representations for Robot ManipulationFoundation Models as World Models: A Foundational Study in Text-Based GridWorldsFrom 2D to 3D Cognition: A Brief Survey of General World ModelsFrom Curiosity to Competence: How World Models Interact with the Dynamics of ExplorationFrom Forecasting to Planning: Policy World Model for Collaborative State-Action PredictionFrom Imitation to Exploration: End-to-end Autonomous Driving based on World ModelFrom Masks to Worlds: A Hitchhiker's Guide to World ModelsFrom Word to World: Can Large Language Models be Implicit Text-based World Models?FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous DrivingGAF: Gaussian Action Field as a Dynamic World Model for Robotic ManipulationGAIA-2: A Controllable Multi-View Generative World Model for Autonomous DrivingGAWM: Global-Aware World Model for Multi-Agent Reinforcement LearningGEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition ControlGLAM: Global-Local Variation Awareness in Mamba-based World ModelGWM: Towards Scalable Gaussian World Models for Robotic ManipulationGaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal GenerationGaussianWorld: Gaussian World Model for Streaming 3D Occupancy PredictionGeneral agents need world modelsGenerating Symbolic World Models via Test-time Scaling of Large Language ModelsGenerative World Modelling for Humanoids: 1X World Model Challenge Technical ReportGenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video GenerationGeoDrive: 3D Geometry-Informed Driving World Model with Precise Action ControlGigaBrain-0: A World Model-Powered Vision-Language-Action ModelGigaWorld-0: World Models as Data Engine to Empower Embodied AIGraph World ModelGrndCtrl: Grounding World Models via Self-Supervised Reward AlignmentHERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and GenerationHERO: Hierarchical Extrapolation and Refresh for Efficient World ModelsHigher Embedding Dimension Creates a Stronger World Model for a Simple Sorting TaskHow Far Are Surgeons from Surgical World Models? A Pilot Study on Zero-shot Surgical Video Generation with Expert AssessmentHow Hard is it to Confuse a World Model?Hunyuan-GameCraft-2: Instruction-following Interactive Game World ModelI2 -World: Intra-Inter Tokenization for Efficient Dynamic 4D Scene ForecastingIPR-1: Interactive Physical ReasonerImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous DrivingInDRiVE: Reward-Free World-Model Pretraining for Autonomous Driving via Latent DisagreementInter-environmental world modeling for continuous and compositional dynamicsInternal World Models as Imagination Networks in Cognitive AgentsJEDI: Latent End-to-end Diffusion Mitigates Agent-Human Performance Asymmetry in Model-Based Reinforcement LearningKAN-Dreamer: Benchmarking Kolmogorov-Arnold Networks as Function Approximators in World ModelsKeyWorld: Key Frame Reasoning Enables Effective and Efficient World ModelsLLM world models are mental: Output layer evidence of brittle world model use in LLM mechanical reasoningLLM-JEPA: Large Language Models Meet Joint Embedding Predictive ArchitecturesLLM-as-a-Judge: Toward World Models for Slate Recommendation SystemsLS-Imagine: Open-World Reinforcement Learning over Long Short-Term ImaginationLUMOS: Language-Conditioned Imitation Learning with World ModelsLaGen: Towards Autoregressive LiDAR Scene GenerationLanguage-Driven Hierarchical Task Structures as Explicit World Models for Multi-Agent LearningLanguage-conditioned world model improves policy generalization by reading environmental descriptionsLarge Emotional World ModelLatent Action World Models for Control with Unlabeled TrajectoriesLatent Chain-of-Thought World Modeling for End-to-End DrivingLatent Policy Steering with Embodiment-Agnostic Pretrained World ModelsLatent-Space Autoregressive World Model for Efficient and Robust Image-Goal NavigationLatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World GenerationLearning Abstract World Models with a Group-Structured Latent SpaceLearning Primitive Embodied World Models: Towards Scalable Robotic LearningLearning Real-World Action-Video Dynamics with Heterogeneous Masked AutoregressionLearning Robot Manipulation from Audio World ModelsLearning To Explore With Predictive World Model Via Self-Supervised LearningLearning World Models for Interactive Video GenerationLearning an Adversarial World Model for Automated Curriculum Generation in MARLLearning to Generate 4D LiDAR SequencesLiDARCrafter: Dynamic 4D World Modeling from LiDAR SequencesLiSTAR: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous DrivingLong-Context State-Space Video World ModelsLongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World ModelLongScape: Advancing Long-Horizon Embodied World Models with Context-Aware MoELongVie 2: Multimodal Controllable Ultra-Long Video World ModelM^3 : A Modular World Model over Streams of TokensMagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive ControlManiGaussian++: General Robotic Bimanual Manipulation with Hierarchical Gaussian World ModelManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual GuidanceMartian World Models: Controllable Video Synthesis with Physically Accurate 3D ReconstructionsMaskGWM: A Generalizable Driving World Model with Video Mask ReconstructionMatrix-Game 2.0: An Open-Source, Real-Time, and Streaming Interactive World ModelMeasuring (a Sufficient) World Model in LLMs: A Variance Decomposition FrameworkMemory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on MinecraftMeta-Reinforcement Learning with Discrete World Models for Adaptive Load BalancingMiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous DrivingMinD: Unified Visual Imagination and Control via Hierarchical World ModelsMindDrive: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous DrivingMindJourney: Test-Time Scaling with World Models for Spatial ReasoningMineWorld: a Real-Time and Open-Source Interactive World Model on MinecraftMissing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image RetrievalMoVieDrive: Multi-Modal Multi-View Urban Scene Video GenerationMoWM: Mixture-of-World-Models for Embodied Planning via Latent-to-Pixel Feature ModulationMobiWorld: World Models for Mobile Wireless NetworkMorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World SimulatorMotus: A Unified Latent Action World ModelMultimodal Dreaming: A Global Workspace Approach to World Model-Based Reinforcement LearningNORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference RewardsNRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World ModelsNatural Building Blocks for Structured World Models: Theory, Evidence, and ScalingNavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation PredictionNavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous EnvironmentsNavigation World ModelsNeural Motion Simulator: Pushing the Limit of World Models in Reinforcement LearningOccProphet: Pushing Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with Observer-Forecaster-Refiner FrameworkOccTENS: 3D Occupancy World Model via Temporal Next-Scale PredictionOccupancy World Model for RobotsOffline Robotic World Model: Learning Robotic Policies without a Physics SimulatorOmniNWM: Omniscient Driving Navigation World ModelsOmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World ModelingOne Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided ExplorationOne Model for All Tasks: Leveraging Efficient World Models in Multi-Task PlanningOpenTwinMap: An Open-Source Digital Twin Generator for Urban Autonomous DrivingOrbis: Overcoming Challenges of Long-Horizon Prediction in Driving World ModelsOther Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latant SpacePIGDreamer: Privileged Information Guided World Models for Safe Partially Observable Reinforcement LearningPIN-WM: Learning Physics-INformed World Models for Non-Prehensile ManipulationParticleFormer: A 3D Point Cloud World Model for Multi-Object, Multi-Material Robotic ManipulationPhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration SynthesisPhysicalAgent: Towards General Cognitive Robotics with Foundation World ModelsPlanning with Reasoning using Vision Language World ModelPlayerOne: Egocentric World SimulatorPragWorld: A Benchmark Evaluating LLMs' Local World Model under Minimal Linguistic Alterations and Conversational DynamicsPrismatic World Model: Learning Compositional Dynamics for Planning in Hybrid SystemsProTerrain: Probabilistic Physics-Informed Rough Terrain World ModelingProphetDWM: ProphetDWM: A Driving World Model for Rolling Out Future Actions and VideosR-WoM: Retrieval-augmented World Model For Computer-use AgentsRELIC: Interactive Video World Model with Long-Horizon MemoryRLVR-World: Training World Models with Reinforcement LearningRadarGen: Automotive Radar Point Cloud Generation from CamerasRaw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)ReSim: Reliable World Simulation for Autonomous DrivingReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online RestorationReimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive ControlRemote Sensing-Oriented World ModelRethinking Driving World Model as Synthetic Data Generator for Perception TasksRethinking the Simulation vs. Rendering Dichotomy: No Free Lunch in Spatial World ModellingRoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic ManipulationRoboScape-R: Unified Reward-Observation World Models for Generalizable Robotics Training via RLRoboScape: Physics-informed Embodied World ModelRobotic World Model: A Neural Network Simulator for Robust Policy Optimization in RoboticsRynnVLA-002: A Unified Vision-Language-Action and World ModelSAMPO: Scale-wise Autoregression with Motion PrOmpt for generative world modelsSCMA: Self-Consistent Model-based Adaptation for Visual Reinforcement LearningSTAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene SimulationSTORM: Search-Guided Generative World Models for Robotic ManipulationSafe Planning and Policy Optimization via World Model LearningScalable Policy Evaluation with Video World ModelsScaling Up Occupancy-centric Driving Scene Generation: Dataset and MethodSceneDiffuser++: City-Scale Traffic Simulation via a Generative World ModelSeeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving SimulationSekai: A Video Dataset towards World ExplorationSemantic Communications with World ModelsSemantic World ModelsSemi-SD: Semi-Supervised Metric Depth Estimation via Surrounding Cameras for Autonomous DrivingSemi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous DrivingSimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World ModelSimple, Good, Fast: Self-Supervised World Models Free of BaggageSimuRA: Towards General Goal-Oriented Agent via Simulative Reasoning Architecture with LLM-Based World ModelSimulating Before Planning: Constructing Intrinsic User World Model for User-Tailored Dialogue Policy PlanningSimulating the Visual World with Artificial Intelligence: A RoadmapSmallWorlds: Assessing Dynamics Understanding of World Models in Isolated EnvironmentsSocial World Model-Augmented Mechanism Design Policy LearningSocial World ModelsSparse Imagination for Efficient Visual World Model PlanningSparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World ModelSparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic QueriesSpatiotemporal Forecasting as Planning: A Model-Based Reinforcement Learning Approach with Generative World ModelsSpeech World Model: Causal State-Action Planning with Explicit Reasoning for SpeechStateSpaceDiffuser: Bringing Long Context to Diffusion World ModelsSurfer: A World Model-Based Framework for Vision-Language Robot ManipulationSynthesizing world models for bilevel planningTeleWorld: Towards Dynamic Multimodal Synthesis with a 4D World ModelTemporal Triplane Transformers as Occupancy World ModelsTerra: Explorable Native 3D World Model with Point LatentsTesserAct: Learning 4D Embodied World ModelsText2World: Benchmarking Large Language Models for Symbolic World Model GenerationThe Double Life of Code World Models: Provably Unmasking Malicious Behavior Through Execution TracesThe Role of World Models in Shaping Autonomous Driving: A Comprehensive SurveyThe Safety Challenge of World Models for Embodied AI Agents: A ReviewThe brain-AI convergence: Predictive and generative world models for general-purpose computationThink Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous VehiclesThinking Ahead: Foresight Intelligence in MLLMs and World ModelsThinking by Doing: Building Efficient World Model Reasoning in LLMs via Multi-turn InteractionTime-Aware World Model for Adaptive Prediction and ControlToward Memory-Aided World Models: Benchmarking via Spatial ConsistencyToward Stable World Models: Measuring and Addressing World Instability in Generative EnvironmentsTowards High-Consistency Embodied World Model with Multi-View Trajectory VideosTowards foundational LiDAR world models with efficient latent flow matchingTraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment VideosTransDreamerV3: Implanting Transformer In DreamerV3Transformer World Model for Sample Efficient Multi-Agent Reinforcement LearningU4D: Uncertainty-Aware 4D World Modeling from LiDAR SequencesUP-VLA: A Unified Understanding and Prediction Model for Embodied AgentUniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous DrivingUnified Vision-Language-Action ModelUnified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic DatasetsUnified World Models: Memory-Augmented Planning and Foresight for Visual NavigationUnlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution ApproachV-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and PlanningVAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM AgentsVCWorld: A Biological World Model for Virtual Cell SimulationVDAWorld: World Modelling via VLM-Directed Abstraction and SimulationVFMF: World Modeling by Forecasting Vision Foundation Model FeaturesVISTAv2: World Imagination for Indoor Vision-and-Language NavigationVL-SAFE: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous DrivingVaViM and VaVAM: Autonomous Driving through Video Generative ModelingVector Quantization in the Brain: Grid-like Codes in World ModelsVehicle Dynamics Embedded World Models for Autonomous DrivingViPRA: Video Prediction for Robot ActionsVid2World: Crafting Video Diffusion Models to Interactive World ModelsVideo World Models with Long-term Spatial MemoryVideoVerse: How Far is Your T2V Generator from a World Model?Vision-Centric 4D Occupancy Forecasting and Planning via Implicit Residual World ModelsVisionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting PlatformVisuomotor Grasping with World Models for Surgical RobotsWMNav: Integrating Vision-Language Models into World Models for Object Goal NavigationWMPO: World Model-based Policy Optimization for Vision-Language-Action ModelsWeb World ModelsWhat Does it Mean for a Neural Network to Learn a "World Model"?What Has a Foundation Model Found? Using Inductive Bias to Probe for World ModelsWhat You Don't Know Can Hurt You: How Well do Latent Safety Filters Understand Partially Observable Safety Constraints?When do Neural Networks Learn World Models?WoMAP: World Models For Embodied Open-Vocabulary Object LocalizationWoW: Towards a World omniscient World model Through Embodied InteractionWorld Model Implanting for Test-time Adaptation of Embodied AgentsWorld Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous DrivingWorld Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task PlanningWorld Models Can Leverage Human Videos for Dexterous ManipulationWorld Models Should Prioritize the Unification of Physical and Social DynamicsWorld Models That Know When They Don't Know: Controllable Video Generation with Calibrated UncertaintyWorld Models Unlock Optimal Foraging Strategies in Reinforcement Learning AgentsWorld Models as Reference Trajectories for Rapid Motor AdaptationWorld Models for Autonomous Navigation of Terrestrial Robots from LIDAR ObservationsWorld Models for Cognitive Agents: Transforming Edge Intelligence in Future NetworksWorld model inspired sarcasm reasoning with large language model agentsWorld-in-World: World Models in a Closed-Loop WorldWorld4Drive: End-to-End Autonomous Driving via Intention-aware Physical Latent World ModelWorld4Omni: A Zero-Shot Framework from Image Generation World Model to Robotic ManipulationWorld4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic ManipulationWorldLens: Full-Spectrum Evaluations of Driving World Models in Real WorldWorldModelBench: Judging Video Generation Models As World ModelsWorldPack: Compressed Memory Improves Spatial Consistency in Video World ModelingWorldPlanner: Monte Carlo Tree Search and MPC with Action-Conditioned Visual World ModelsWorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World ModelingWorldVLA: Towards Autoregressive Action World ModelWristWorld: Generating Wrist-Views via 4D World Models for Robotic ManipulationX-WIN: Building Chest Radiograph World Model via Predictive SensingXray2Xray: World Model from Chest X-rays with Volumetric ContextZero-Splat TeleAssist: A Zero-Shot Pose Estimation Framework for Semantic TeleoperationZero-shot World Models via Search in Memoryseq-JEPA: Autoregressive Predictive Learning of Invariant-Equivariant World Models3D-VLA: A 3D Vision-Language-Action Generative World ModelAD3: Implicit Action is the Key for World Models to Distinguish the Diverse Visual DistractorsAVID: Adapting Video Diffusion Models to World ModelsAdaptive World Models: Learning Behaviors by Latent Imagination Under Non-StationarityAdvancing Humanoid Locomotion: Mastering Challenging Terrains with Denoising World Model LearningAgent Planning with World Knowledge ModelAn Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted TrainingBEVWorld: A Multimodal World Model for Autonomous Driving via Unified BEV Latent SpaceBWArea Model: Learning World Model, Inverse Dynamics, and Policy for Controllable Language GenerationBounded Exploration with World Model Uncertainty in Soft Actor-Critic Reinforcement Learning AlgorithmCarDreamer: Open-Source Learning Platform for World Model based Autonomous DrivingCarFormer: Self-Driving with Learned Object-Centric RepresentationsCausal World Representation in the GPT ModelCityBench: Evaluating the Capabilities of Large Language Model as World ModelCoDreamer: Communication-Based Decentralised World ModelsCognitive Map for Language Models: Optimal Planning via Verbally Representing the World ModelCognitively Inspired Energy-Based World ModelsCompete and Compose: Learning Independent Mechanisms for Modular World ModelsCopilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete DiffusionDINO-WM: World Models on Pre-trained Visual Features enable Zero-shot PlanningDOME: Taming Diffusion Model into High-Fidelity Controllable Occupancy World ModelDexSim2Real$^2$: Building Explicit World Model for Precise Articulated Object Dexterous ManipulationDiffusion for World Modeling: Visual Details Matter in AtariDo Transformer World Models Give Better Policy Gradients?Doe-1: Closed-Loop Autonomous Driving with Large World ModelDream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with ImaginationDreamSmooth: Improving Model-based Reinforcement Learning via Reward SmoothingDreaming of Many Worlds: Learning Contextual World Models Aids Zero-Shot GeneralizationDriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video GenerationDriveDreamer: Towards Real-world-driven World Models for Autonomous DrivingDriveGenVLM: Real-world Video Generation for Vision Language Model based Autonomous DrivingDriveWorld: 4D Pre-trained Scene Understanding via World Models for Autonomous DrivingDriving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous DrivingDriving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous DrivingDrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World ModelDrivingWorld: Constructing World Model for Autonomous Driving via Video GPTEVA: An Embodied World Model for Future Video AnticipationEfficient Exploration and Discriminative World Model Learning with an Object-Centric AbstractionEfficient World Models with Context-Aware TokenizationEmergence of Implicit World Models from Mortal AgentsEnhancing End-to-End Autonomous Driving with Latent World ModelEvaluating the World Model Implicit in a Generative ModelExploring the Interplay Between Video Generation and World Models in Autonomous Driving: A SurveyGeneralized Predictive Model for Autonomous DrivingGenerative Emergent Communication: Large Language Model is a Collective World ModelGenie: Generative Interactive EnvironmentsGrounded Answers for Multi-agent Decision-making Problem through Generative World ModelGrounding Large Language Models In Embodied Environment With Imperfect World ModelsHarmonyDream: Task Harmonization Inside World ModelsHierarchical World Models as Visual Whole-Body Humanoid ControllersHieros: Hierarchical Imagination on Structured State Space Sequence World ModelsHoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous DrivingHow Far is Video Generation from World Model: A Physical Law PerspectiveIGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AIImagine-2-Drive: High-Fidelity World Modeling in CARLA for Autonomous VehiclesImproving Token-Based World Models with Parallel Observation PredictionInfinityDrive: Breaking Time Limits in Driving World ModelsIs Sora a World Simulator? A Comprehensive Survey on General World Models and BeyondIs Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web AgentsLanguage Agents Meet Causality -- Bridging LLMs and Causal World ModelsLearning Latent Dynamic Robust Representations for World ModelsLearning Multiple Probabilistic Decisions from Latent World Model in Autonomous DrivingLearning World Models for Unconstrained Goal NavigationLearning and Leveraging World Models in Visual Representation LearningLidarDM: Generative LiDAR Simulation in a Generated WorldMAMBA: an Effective World Model Approach for Meta-Reinforcement LearningMagicDrive: Street View Generation with Diverse 3D Geometry ControlMagicTime: Time-lapse Video Generation Models as Metamorphic SimulatorsMaking Large Language Models into World Models with Precondition and Effect KnowledgeMaking Offline RL Online: Collaborative World Models for Offline Visual Reinforcement LearningManiGaussian: Dynamic Gaussian Splatting for Multi-task Robotic ManipulationMastering Memory Tasks with World ModelsMitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World ModelsMotion Prompting: Controlling Video Generation with Motion TrajectoriesMulti-Task Interactive Robot Fleet Learning with Visual World ModelsMultimodal foundation world models for generalist embodied agentsOccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous DrivingOccSora: 4D Occupancy Generation Models as World Simulators for Autonomous DrivingOccWorld: Learning a 3D Occupancy World Model for Autonomous DrivingOne-shot World Models Using a Transformer Trained on a Synthetic PriorOwl-1: Omni World Model for Consistent Long Video GenerationPIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic ManipulationPWM: Policy Learning with Large World ModelsPhysical Informed Driving World ModelPlanning with Adaptive World Models for Autonomous DrivingPredicting vs. Acting: A Trade-off Between World Modeling & Agent ModelingProbing Multimodal LLMs as World Models for DrivingR-AIF: Solving Sparse-Reward Robotic Tasks from Pixels with Active Inference and World ModelsRenderWorld: World Model with Self-Supervised 3D LabelRepresenting Positional Information in Generative World Models for Object ManipulationReward-free World Models for Online Imitation LearningRoboDreamer: Learning Compositional World Models for Robot ImaginationSafeDreamer: Safe Reinforcement Learning with World ModelsScaling Laws for Pre-training Agents and World ModelsSimuDICE: Offline Policy Optimization Through World Model Updates and DICE EstimationStoryWeaver: A Unified World Model for Knowledge-Enhanced Story Character CustomizationTD-MPC2: Scalable, Robust World Models for Continuous ControlTerra ACT-Bench: Towards Action Controllable World Models for Autonomous DrivingThe Matrix: Infinite-Horizon World Generation with Real-Time Moving ControlThink2Drive: Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous DrivingTokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous DrivingTowards Physically Interpretable World Models: Meaningful Weakly Supervised Representations for Visual Trajectory PredictionTowards Unraveling and Improving Generalization in World ModelsTransformers Use Causal World Models in Maze-Solving TasksTransformers and Slot Encoding for Sample Efficient Physical World ModellingUnO: Unsupervised Occupancy Fields for Perception and ForecastingUnderstanding World or Predicting Future? A Comprehensive Survey of World ModelsUniMLVG: Unified Framework for Multi-view Long Video Generation with Comprehensive Control Capabilities for Autonomous DrivingUnleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video GenerationUrbanWorld: An Urban World Model for 3D City GenerationVidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot ManipulationVista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilityVisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot PlanningWHALE: Towards Generalizable and Scalable World Models for Embodied Decision-makingWeb Agents with World Models: Learning and Leveraging Environment Dynamics in Web NavigationWorld Model on Million-Length Video And Language With RingAttentionWorld Model-based Perception for Visual Legged LocomotionWorld Models Increase Autonomy in Reinforcement LearningWorld Models for Autonomous Driving: An Initial SurveyWorld Models with Hints of Large Language Models for Goal AchievingWorld Models: The Safety PerspectiveWorldDreamer: Towards General World Models for Video Generation via Predicting Masked TokensADriver-I: A General World Model for Autonomous DrivingCategorical Traffic Transformer: Interpretable and Diverse Behavior Prediction with Tokenized LatentFOCUS: Object-Centric World Models for Robotics ManipulationGAIA-1: A Generative World Model for Autonomous DrivingLearning to Model the World with LanguageMUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric RepresentationsSTORM: Efficient Stochastic Transformer based World Models for Reinforcement LearningTask Aware Dreamer for Task Generalization in Reinforcement LearningTrafficBots: Towards World Models for Autonomous Driving Simulation and Motion PredictionTransformer-based World Models Are Happy with 100k InteractionsTransformers are Sample Efficient World ModelsUniWorld: Autonomous Driving Pre-training via World ModelsDayDreamer: World Models for Physical Robot LearningDeep Hierarchical Planning from PixelsDreamerPro: Reconstruction-Free Model-Based Reinforcement Learning with Prototypical RepresentationsDreamingV2: Reinforcement Learning with Discrete World Models without ReconstructionHierarchical Model-Based Imitation Learning for Planning in Autonomous DrivingIso-Dream: Isolating and Leveraging Noncontrollable Visual Dynamics in World ModelsSymphony: Learning Realistic and Diverse Agents for Autonomous Driving SimulationMastering Atari with Discrete World ModelsDream to Control: Learning Behaviors by Latent ImaginationPlanning to Explore via Self-Supervised World ModelsWorld ModelsPerceptUI: PerceptUI: LLM Agents as Human-Aligned Synthetic Users for UI/UX EvaluationMemoBench: Benchmarking World Modeling in Dynamically Changing EnvironmentsEO-WM: A Physically Informed World Model for Probabilistic Earth Observation ForecastingHallucination in World Models is Predictable and PreventablePhysiFormer: Learning to Simulate Mechanics in World SpaceFast LeWorldModelPredRNN: A Recurrent Neural Network for Spatiotemporal Predictive LearningSimVP: Simpler yet Better Video PredictionDiffusion Models for Video Prediction and InfillingTemporal Attention Unit: Towards Efficient Spatiotemporal Predictive LearningOpenSTL: A Comprehensive Benchmark of Spatio-Temporal Predictive LearningMCVD: Masked Conditional Video Diffusion for Prediction, Generation, and InterpolationMemLearner: Learning to Query Context memory for Video World ModelsLearning Transferable Dynamics Priors from Action to World ModelingDreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World ModelValdi: Value Diffusion World ModelsABot-M0.5: Unified Mobility-and-Manipulation World Action ModelWorldDirector: Building Controllable World Simulators with Persistent Dynamic MemoryWorldOdysseyBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World ModelsDVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic ManipulationFrom World Models to World Action Models: A Concise Tutorial for RoboticsOPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive ExplorationCertified World Models as Sensing Clocks: Drift-Aware Deadlines for Active PerceptionSafe and Adaptive Cloud Healing: Verifying LLM-Generated Recovery Plans with a Neural-Symbolic World ModelPredicting Closed-Loop Performance of Latent World Models: Offline Checkpoint Selection for MPC and Model-Based RL Under Non-Markovian Rewards in LunarLanderPhysMani: Physics-principled 3D World Model for Dynamic Object ManipulationPWM-ArtGen: Part World Model for Articulated Object GenerationBridge-WA: Predicting Where and How the World Changes for Robotic ActionACID: Action Consistency via Inverse Dynamics for Planning with World ModelsWorldSample: Closed-loop Real-robot RL with World ModellingOWMDrive: Causality-Aware End-to-End Autonomous Driving via 4D Occupancy World ModelSelf-Evolving World Models for LLM Agent PlanningLong-term Traffic Simulation via Structured Autoregressive ModelingDelta-JEPA: Learning Action-Sensitive World Models via Latent Difference DecodingOne Video, One World: Turning Monocular Video into Physical 4D ScenesWorld-Model Collapse as a Phase TransitionAsk the World Before Acting: Budgeted Environment Probing for World-Model CalibrationWorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World ModelsAdaJEPA: An Adaptive Latent World Model3D Point World Models: Point Completion Enables More Accurate Dynamics LearningRetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in RetailMulti-scale Mixture of World Models for Embodied Agents in Evolving EnvironmentsPath Planning in Physically Viable World ModelsRoboWorld: Fast and Reliable Neural Simulators for Generalist Robot Policy EvaluationIRASim: A Fine-Grained World Model for Robot ManipulationUnified Video Action ModelEnerVerse-AC: Envisioning Embodied Environments with Action ConditionVLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World SimulatorsGigaWorld-1: A Roadmap to Build World Models for Robot Policy EvaluationWorldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action ControlDynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual DynamicsLast-Meter Precision Navigation for UAVs: A Diffusion-Refined Aerial Visual Servoing ApproachLearning Task-Sufficient World Models by Synergizing Agentic Exploration and Structured ModelingOperator-on-F complements value-equivalence: a planning-time diagnostic for latent world modelsGeographic Diversity Beats Data Volume for Cross-Domain Generalization in Zero-Label JEPA Driving World ModelsMask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World ModelsKAM-WM: Kinematic Affordance Maps from Latent World Models for Robot ManipulationDSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot ManipulationQantara: Bridge-Flow Training for Multi-Paradigm JEPA ControlMoP-JEPA: Hard-Assigned Predictor Mixtures for Stochastic JEPA World ModelsMultiplayer Interactive World Models with Representation AutoencodersDeform360: A Massive Multi-view Visuotactile Dataset for Deformable World ModelsLongCat-Video Technical ReportRynnWorld-4D: 4D Embodied World Models for Robotic ManipulationRynnWorld-Teleop: An Action-Conditioned World Model for Digital TeleoperationAlayaWorld: Long-Horizon and Playable Video World GenerationImagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model FailureWildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial IntelligenceInfinite Worlds with Versatile InteractionsPanoWorld: Real-World Panoramic GenerationFlow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic SimulationPathdreamer: A World Model for Indoor NavigationV-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation ModelsXiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation ModelImproving Multi-Step Prediction of Learned Time Series ModelsVersatile Behavior Diffusion for Generalized Traffic Agent SimulationBehaviorGPT: Smart Agent Simulation for Autonomous Driving with Next-Patch PredictionLearning to drive from a world on railsSelf in Space: Benchmarking Self-Awareness and Spatial Cognition in UAV Embodied IntelligenceGigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearchCollision Avoidance Detour for Multi-Agent Trajectory ForecastingHierarchical Denoising For Multi-Step Visual ReasoningRxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and ImaginationFrom Pixels to States: Rethinking Interactive World Models as Game EnginesCan World Models Benefit VLMs for World Dynamics?LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data IngestionRethinking Video Generation Model for the Embodied WorldMixture of Contexts for Long Video GenerationRecurrent Autoregressive Diffusion: Global Memory Meets Local AttentionE3DGS: Unified Geometric-Photometric Equivariance for 3D Gaussian Splatting via Color-as-Geometry EmbeddingSeerGuard: A Safety Framework for Mobile GUI Agents via World Model PredictionOrbis 2: A Hierarchical World Model for DrivingDSWorld: A Data Science World Model for Efficient Autonomous AgentsLearning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RLPAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action ExplorationGeoWorldAD: Geometry World Action Model for Autonomous DrivingThinking in Video: Can Video Generators Really Reason About the Real World?Reinforcement Learning: From Algorithms To Foundation ModelsPlanning with Transformers: Chain of Computation and Structured Context WindowsMobile Network Control with a World ModelSAGE: Subgoal-Conditioned Action Generation for Latent World Model PlanningMasked Visual Actions for Unified World ModelingMasked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RLOpen-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied LearningApple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical IntelligenceEvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary WorldAlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical ReportGenerative World Renderer at the Speed of PlayABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPUStreaming Multi-Agent Autoregressive Diffusion Model with World State RegistersStable Video Infinity: Infinite-Length Video Generation with Error RecyclingGameFactorly: Creating New Games with Generative Interactive VideosLoViC: Efficient Long Video Generation with Context CompressionPlayable Video GenerationGenerative Pre-trained Autoregressive Diffusion TransformerVideo-GPT via Next Clip DiffusionRoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic ManipulationVideo Generation Models in Robotics - Applications, Research Challenges, Future DirectionsWonder: Video World Model Done BetterWorldDiT: A Unified Diffusion Architecture for World and Action ModelingTemporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive ControlVisualPatchWorld: Code World Models as Latent Structured Representations for PlanningACE-Data-0: Human-Centric Ambient Capture as Embodied Data EnginePhiZero: A World Model Built Around Physical LanguageStatePlay: State-Aware Game World Models for Mechanics-Consistent GenerationINTACT: Isomorphic Intent-to-Action Learning for Search-Free World ModelsShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its ShadowMultiWorld: Scalable Multi-Agent Multi-View Video World ModelsPlanning from Pixels using Inverse Dynamics ModelsSlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video GenerationQQWorld: Quantile-Quantile Matching for World Model RegularizationWorld Action Planner: Generalizable Decision-Making with Action-Conditioned World ModelsODEWorld: A Continuous Predictive Architecture via Physical-Time FlowSG-WAM: Self-Guided World Modeling in Geometry-Aware Policy SpaceHelloWorld: Enabling Socially Interactive Characters in Video World ModelsMiniWorld: Democratizing the Training of Video World Models from ScratchQuo Vadis, World Modeling?WorldClaw: Agentic 3D Open-World Generation at ScaleWorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World ModelsMirage 2: Ai-native ugc game engine powered by real-time world modelsScaling Agent Learning via Experience Synthesis
Resources
Tags
size_categories:1K<n<10Kmodality:imagemodality:videolibrary:datasetslibrary:mlcroissantregion:us